Claude Opus 4.8 Update: What Changed?
45sTech enthusiasts crave quick summaries of major AI updates, and this intro promises a direct comparison with GPT 5.5.
▶ Play ClipThe video covers the release of Claude Opus 4.8 by Anthropic, detailing its improvements, pricing, and practical comparisons with GPT 5.5. The presenter demonstrates the model's enhanced performance in autonomous tasks, project completion, and honesty, while also introducing new features like dynamic workflows and effort levels.
Anthropic has released Opus 4.8, an update to their AI model. The video aims to show changes, impact on usage, and comparisons with other models.
Benchmarks are standardized tests comparing AI performance. Opus 4.8 scored highest in autonomous computer use tasks, surpassing GPT 5.5.
Opus 4.8 is the only model to complete 100% of projects in the Super Agent Benchmark, indicating high reliability for end-to-end tasks.
Opus 4.8 is four times less likely to deliver incorrect results without verification, reducing hallucinations and false completions.
A new feature that allows the model to spawn multiple sub-agents for complex tasks, available only on the Max plan. It consumes more tokens.
Users can now choose from low, medium, high, extra, and maximum effort levels, controlling how much the model 'thinks' on a task.
Despite being an update, the price remains the same as previous versions, still one of the most expensive AI models.
Anthropic hinted at a future Opus 4.9 model, currently being tested with specific companies for safety before wider release.
The presenter tested a logic question: 'I live 5 km from the car wash. Should I go by car or walk?' Opus 4.8 correctly answered 'by car', unlike previous versions.
Both Opus 4.8 and GPT 5.5 created a snake game in HTML. Opus 4.8's version was playable and better executed.
Opus 4.8 and GPT 5.5 both created interactive particle animations. GPT's version was visually preferred by the presenter.
Opus 4.8 produced a more functional drawing app with brush, eraser, and spray paint, outperforming GPT 5.5's version.
Claude Opus 4.8 brings significant improvements in honesty, reliability, and task completion, with new features like dynamic workflows and effort levels. While it remains expensive, its performance in coding and reasoning tests shows it is a strong competitor to GPT 5.5.
"The title accurately reflects the update and comparison, though it slightly overpromises by implying a full tutorial on usage."
What is the name of the new Claude model released by Anthropic?
Opus 4.8
00:01
In which benchmark did Opus 4.8 score higher than GPT 5.5?
Autonomous computer use tasks
02:27
What percentage of projects did Opus 4.8 complete in the Super Agent Benchmark?
100%
03:54
How much more honest is Opus 4.8 compared to previous versions?
Four times more honest
05:02
What is the name of the new feature that allows Opus 4.8 to spawn multiple sub-agents?
Dynamic Workflow
06:10
Which plan is required to access the Dynamic Workflow feature?
Max plan
06:10
How many effort levels are available in Opus 4.8?
Five: low, medium, high, extra, and maximum
07:05
Did the price of Opus 4.8 increase compared to previous versions?
No, the price remains the same.
08:05
What is the cost for output of 1 million tokens in Opus 4.8?
$5
08:17
Which model performed better in the drawing app test according to the presenter?
Opus 4.8
16:39
100% Project Completion
Opus 4.8 is the only model to complete all projects in the Super Agent Benchmark, demonstrating exceptional reliability.
03:39Four Times More Honest
A significant reduction in false completions, making the model more trustworthy for critical tasks.
05:02Dynamic Workflow
A novel feature that enables complex multi-agent task execution, though limited to the Max plan.
06:10Price Stability
Unlike typical updates, the price remains unchanged, offering better value for existing users.
08:05Improved Reasoning
Opus 4.8 correctly answers a logic question that previous versions failed, showing enhanced reasoning capabilities.
11:46[00:01] on the market has just received an update. Opus has now become Opus update. Opus has now become Opus 4.8. H Tropic just released this update. I'm going to show you what changes, how it might affect the
[00:14] way you use the tool, whether it will improve anything for you, whether it will become more expensive, cheaper, or faster. In this video, I want to share this information and show some comparisons so you can
[00:26] understand and know how to apply it in your daily life, whether for your 'm Rob. Here we talk about canvas, artificial intelligence, and tools that can help you be much more productive and make money online.
[00:39] subscribe to the channel now. We have videos like this every week, okay? And if you use the cloud in your daily life, I'm sure this will help you and keep you updated, and you'll get ahead that way, okay?
[00:54] Let's talk about these modifications. Whenever there's news about the cloud, especially here, there's a huge hype on the internet. People talk a lot about the model, and in fact it is the smartest and most widespread model
[01:09] in terms of productivity, the first to bring several things that other models didn't have. So, it's truly fantastic, but it's also one of the most expensive artificial intelligence models on the market. And that causes many
[01:22] people not to use it as much, they use the cloud together with other models, you know, other artificial intelligences to complement it. But before we talk about what has changed, we need to understand how these tests are conducted.
[01:34] talking about regarding the numbers. If you go to the Antropic website, you'll see all these numbers here. This is in English, obviously. There are several posts here, another post is linked. You go there and see what's
[01:48] inside. This is like a general summary. But what do these numbers here represent ? They are benchmarks, meaning they are tests that are done to bring us these results, but often it's not very clear what they
[02:00] ENEM exam, a test that these artificial intelligences administer. Everyone takes the same test, and then they compare the results to see who performs best at this or that task. In this case, these are real tasks; it's a
[02:13] comparison based on the Antropic documentation, and that way they can collect information to see the practical impact of the tool. And that 's one of the reasons why people like this model so much, because it
[02:27] tends to be very honest about what it delivers. One of the comparisons is to have artificial intelligence use the internet, use the computer on its own. In other words, an example here is if you have the cloud extension on your
[02:40] and ask it to do a survey for you, ask it to fill out a form, something along those lines. In this case, the cloud scored 4.8, much higher than the other models, okay? I didn't put the exact value of GPT 5.5 here because there's a
[02:55] value there that, if you look in the documentation—I read the documentation at the value is the exact value, the number is the exact number. So I decided not to include it here for that very reason, but it scored much higher than the other
[03:09] models that already exist. So, in other words, for autonomous tasks, it remains one of the best and is now even above the GPT 5.5, which was the best until now. Apparently, Open is already working
[03:24] on the release of GPT 5.6. It might be out by the end of this video, who knows? Another thing is in relation to completing a project. So, there's a benchmark, it's called the Super Agent Benchmark. This benchmarking they do
[03:39] is designed to make artificial intelligence complete a project from start to finish. So they hand this project over to the model, and the model needs to complete the entire project. And the only model that completed 100% of all
[03:54] only model that completed 100% of all projects was the Opus 4.8. That's according to Antropic itself. And how does this reflect in your daily life? Well, you can trust the model more in theory, right? So you can give him a task and he will complete
[04:07] that task. Many people were complaining about Opus 4.7, I don't know about you, but I use it a lot in my daily life and I really felt that there were some things it didn't deliver, but now apparently it will deliver much more. And
[04:19] I saw. Let's see how it performs in my day-to-day work. Correct answers. So they have a benchmark where they have models analyze contracts and other things, and respond to see if the answers are actually
[04:33] correct or not. And the one that came in first place was Opus. Obviously, he was ahead of the others here, presenting this information correctly and getting everything right. So it's for huge tasks like analyzing contracts,
[04:49] checking invoices, things where the model really needs to be very reliable. So, in this way, they are marketing the model so that tools can be created or people can use this model for other things
[05:02] related to legal issues, medicine, and other things. Now, one actually an improvement that Propropic mentioned, is that the model is four times more honest. Well, I would expect the model to be honest, right? All
[05:18] apparently it's four times more honest. What does this mean here so you can understand? Sometimes artificial intelligence, in order to please you, is something that GPT always does and continues to do, but not
[05:31] the codex models, the models geared towards coding, okay? But the thing is, it wasn't a hallucination at all. He would say he finished a task, when in fact he didn't; he didn't. So, for example, create
[05:45] a post about such and such. He would bring the information, "Look, it's here, it turned out well," but then you go there and think, "No, it didn't turn out well, this isn't what I wanted, right?" Now, Opus 4. It will bring this information, right, in a much more precise way. So
[05:58] it's four times less likely, meaning you have four times less chance of the model delivering times less chance of the model delivering something wrong without first checking for it. Finally, it became more reliable. And they launched something called Dynamic
[06:10] Workflow. This is super interesting, but from what I've seen it's only available for certain plans. Here I have the Max plan, for example, it's don't have the Max plan, which is the most expensive plan in the Cloud, you won't be able to
[06:24] access this. This dynamic workflow is basically as follows. Before, you would task. He could even execute sub-agents before, and now you give him a task and he'll execute multiple agents. So it's as if he puts several other
[06:38] models there to work for you, several other little robots, so to speak, perform this task. Obviously, this will consume more of your tokens while you're performing that task. So you have to be very
[06:53] careful when using this. I myself don't know if I'll use it, because anyone who uses the cloud knows that the tokens disappear really quickly, right? And now they've also released five levels of effort. You have low, medium,
[07:05] high, extra, and maximum levels. So now when you access the cloud here, you'll see that version 4.8 is already available. And here at the bottom, look, you have the effort levels: high,
[07:22] medium, low, extra, and max. Most likely, at Max, you'll consume a lot of your tokens in your daily use. But this is really good when you're solving a problem, creating something
[07:37] super complex where you need the intelligence and the model to really think harder, check the information more, so you give it a certain effort, like a turbo boost, to help. And the fast mode for those who
[07:50] code, for those who create, it got faster, it got quicker. I myself don't also become cheaper than the previous version, and the price remains the same, meaning it hasn't increased. This is one of the first times I've seen them release a
[08:05] new model, an update, but not increase the price. It remains the most expensive artificial intelligence model, one of the most expensive on the market. So today, for you to input, that is, for you to put in 1 million tokens—each token
[08:17] would be almost like a word—it costs [amount], and for you to have the output, that is, for the artificial intelligence to bring you back 1 million tokens, right? In other words, what million tokens, right? In other words, what she delivers costs $5. I
[08:32] mean, that's a lot. And here are several comparisons between Cloud, GPT, and Gemeni. Gemini is falling quite a bit behind the others, isn't it? Well, Cloud is scoring very well here, except for the benchmark
[08:48] for terminal usage where GPT 5.5 continues to score slightly better. before we went there to do some testing? 84% better at using the computer
[09:00] independently. He was the only agent who completed that test, the super beginning to end. There he became four times more honest. Now you have five effort levels and the price hasn't changed at all. And what Antropic said, including
[09:15] in the post they have on their website, is that very soon they will be Mitos, right? We were already expecting this model; it's being used by some specific companies because it 's an extremely powerful model.
[09:30] sure that this model won't be used for malicious purposes. So, they want to ensure they have mechanisms in place to keep it safe. I believe
[09:42] to kind of market this model right now, right? Launching this entry-level model is like giving it a facelift. from the cloud. If you go to Twitter, you'll see that there are several tests that people
[09:57] are doing, comparing the tools here. Some show Opus 4.7 as tools here. Some show Opus 4.7 as better, others show 4.8 as better. This guy here created something like a Minecraft game, spending money creating
[10:11] this visual aspect using Opus. And here's another test that's very common, people do it quite often, which is asking how many fingers are on this hand. Then he posted a picture here with an extra finger, like
[10:24] a thumb. And the artificial intelligence responded here that there were indeed four fingers there, a thumb and an extra thumb. It's in English, it would be four fingers, one thumb, right, and an extra thumb, that's it. And that makes it very clear
[10:39] here, and in several other tests that people have done. Well, if you use the have the application installed on your computer so you can use it. If you different than using C there. If you use Coworking, I really recommend
[10:53] that you use the app, because it's the only way, including computer. I have the Max plan here. When you access it, it's already available here in the chat version, it's already here inside the coworking space as well.
[11:06] this presentation. Well, I brought the information and he gave a presentation here for me. Here you can gain access. In this case, I'm going to ask for a new task. And then you'll see that you can already
[11:18] access it. Look, version 4.8 is already here. Selecting the effort level and the models, right? And it's also available here in the cloud . I have
[11:31] several projects here that I've been working on and I'm super excited to be able to take the test. I still have access to the Opus 4.7 model here, it's legacy and get more expensive, will it get cheaper? Well, from what I've seen, the price is the same.
[11:46] , okay? Let's do a reasoning test here so we can see. This is something that many models get wrong or respond to in a very strange way. I respond to in a very strange way. I live only 5 km from the car wash, the
[12:00] closest one to my house. Should I go by car or walk? So here I am on the Opus 4.8 model. Let's see what he'll answer. Oh, go by car or on foot to wash the car. That's a good one. Are you going by car? Of course. A
[12:14] car wash is for washing your car. So there's no point in walking there and leaving your dirty car at home, unless you want to wash your boots. You can walk there . I asked the same question here for GPT 5.5. He also replied: "Go by
[12:30] car." I had asked this question for Opus 4.6 at the time, and also for 4.7. He said, "Walk because it's healthier." So here he has already answered something correct. Here I'm going to give you a prompt, which is as follows: Creates a
[12:43] prompt, which is as follows: Creates a complete snake game in HTML in a single file, dark background, green snake, neon lights, red apple, scoreboard at the top, speed increases every five points. Game over with the option to restart everything in a single
[12:57] HTML file. Okay, I'll ask him to create it for me here, and I'll do the same create it for me here, and I'll do the same thing here within the 5.5 GPT. Oh, GPT already started creating this for us really quickly, didn't it? Look how interesting,
[13:09] already found out that I have a skill, so that's really cool, it's a skill that I have here. So he sought out that skill to figure out how to build the best graphical interface based on what I have
[13:23] that skill, he wouldn't have looked for it. That's why it's taking him a little longer. Hey, the GPT here has already finished and even said: "Just copy and preview here so you can see it. Oh, the preview is already here. I'll ask to play
[13:37] again. Oh, it's working, you see? I hit it there, and it worked. Since I couldn't open the entire screen here, look. The cloud here has also ended. So, let's ask them to begin. They made something very similar, but the cloud version is better here
[13:49] because I can actually play the game. Man, it's crazy to know you can do that today. Look, the speed is really increasing here, is n't it? Game over. That's great, man. Very good indeed. Let's order another one here.
[14:03] Create an interactive HTML particle animation . When the mouse hovers across the screen, colored particles appear and connect with lines if they are nearby. Dark theme, purple and orange color scheme,
[14:18] Dark theme, purple and orange color scheme, unique HTML also, 4.8 resolution. Let's also look at the Most likely, GPT will be faster here because it's not seeing any skills. In this case, Opus went there to be able to look at the skill. He's already read it
[14:32] here, right? Oh, Cloud was faster here, huh? It's over, then. It looks good here. He already created it here for me. The GPT has also ended. Let's see how it turned out. I really liked the GPT here, okay? Honestly, but it remained
[14:46] written here in living particles. I liked it a lot. I think it looks better here on GPT. I don't know. Leave a comment so I can find out. But what does this prove to us, right? It only proves the capability of these models.
[14:58] Look how advanced we are compared to previous models that already existed. They can certainly do it here, but just imagine how quickly it would be possible to quick test here to create an
[15:13] HTML drawing application with a brush and eraser, and we'll do the same thing here inside GPT. Let's go. I started here first with Opus Dei. Oh, Opus is already doing the research over there. GPT was already there, kicking ass and creating things, right? One
[15:30] thing I've always really liked about Antropic's designs is that they don't appeal to you. So, if you ask a question, it asks you questions back to help you reason it out. He's not going to be there all the time telling you what you
[15:44] 're doing so wonderfully, which is something GPT tends to do quite a lot there, So, look, he's going to start creating here for us. It's very likely that GPT has already finished. Let's see. No, not yet. It's almost finished here, look.
[15:57] It ended. GPT was faster, but here he was looking, obviously, at the he already created it for me. Let's see. Oh, I have an eraser, a brush, and spray paint. Hmm, my ex isn't doing too well, is he? I could have put it here. Let's take a look here.
[16:11] This isn't very good, look. See? My mouse is over there on the other side. And then it gave me the option to save as PNG and clear it. Oh, when I click save, it doesn't work either. The rubber band works more or less.
[16:25] Okay, so it didn't turn out very well. Let's see our friend Claudão here. Comment you use the Cloud and want me to create more content about the tool, I can provide it. I use it in my daily life, for my work. Besides the channel, I
[16:39] also work, developing and testing various models in my day-to-day life. Look how cool, Claudão already won here, right? The graphical interface alone has already improved significantly. It's already here. Let's see. Grab some color here, look.
[16:55] Much better, much more interactive, oh. Just look. Much better, right? And here I if I can save it. If I click save here, oh, I can't save, no but now we can understand. Comment below what you
[17:10] think. I wanted to bring this update here for you all. I know there are many people interested in using it. Please also comment on whether you use the cloud or here for you to watch. See you
[17:22] here for you to watch. See you next time. See you later.
⚡ Saved you 0h 17m reading this? Transcribe any YouTube video for free — no signup needed.