---
title: 'AI Bizarro Land: A Week of Major Model Releases'
source: 'https://youtube.com/watch?v=FluKUJyeYD8'
video_id: 'FluKUJyeYD8'
date: 2026-09-04
duration_sec: 446
channel: 'Fireship'
---

# AI Bizarro Land: A Week of Major Model Releases

> Source: [AI Bizarro Land: A Week of Major Model Releases](https://youtube.com/watch?v=FluKUJyeYD8)

## Summary

The video covers a chaotic week in AI, with Anthropic, Meta, and OpenAI all releasing major models. It breaks down each release, highlights key benchmarks and real-world applications, and discusses the messy rollout of OpenAI's GPT-6 Astra.

### Key Points

- **Three Major AI Releases** [00:00] — Anthropic released Fable and Mythos 5.1, Meta released Muse Spark 1.3, and OpenAI announced GPT-6 Astra, creating a whirlwind of AI news.
- **Anthropic's Fable and Mythos 5.1** [01:36] — Fable 5.1 debugged a five-year-old crash by disassembling a vendor library, while Mythos 5.1 improved protein design success from 10% to 50%.
- **Meta's Muse Spark 1.3** [02:31] — Meta's Muse Spark 1.3 offers a contributor tier at $0.10 in and $0.20 out, with a double-digit percentage of developers opting in.
- **OpenAI's GPT-6 Astra Launch** [03:18] — OpenAI's GPT-6 Astra launch was chaotic, with a temporary takedown and apology from Sam Altman. The model wasn't publicly available immediately.
- **Astra's Benchmarks and Capabilities** [04:47] — Astra scored 73% on OS World, 100% on ExportBench, and 99% on ArcAGI3, and is the first model to hit OpenAI's critical cyber threshold.
- **Independent Testing Results** [06:22] — Independent testing gave Astra a 61 on the Intelligence Index, same as GPT-5.6 and 5 points behind Fable 5.1.

### Conclusion

The week showcased rapid AI advancements from Anthropic, Meta, and OpenAI, but also highlighted the gap between marketing hype and independent benchmarks. As models push boundaries, the industry faces both exciting possibilities and messy rollouts.

## Transcript

The famous American philosopher Smash Mouth once said, the years start coming and they don't stop coming. And I've never understood those words more deeply than I did this week after what felt like years in AI Bizarro Land. First, on Tuesday, Anthropic released Fable and Mythos 5.1,
which they're calling the world's most advanced models for coding and knowledge work. Then on Wednesday, Meta released Muse Spark 1.3, which Zuck says is a frontier model that's almost too cheap to meter. Then finally, on Thursday, OpenAI announced GPT-6 Astra,
which President Greg Brockman says is AGI if you're one of the handful of lucky influencers or enterprise customers who got access to it. Unfortunately, I think I've made too many gay scam open jokes over the years to make the cut,
but in this video, I'll do my best to break down the new models, see how they anointed or liken Astra, and look into the totally coincidental timing where the internet broke right as it was released. It is September 4th, 2026, and we're watching the code report.
Normally, the first week of September is a quiet one because all the rich VCs and poor project managers are busy appropriating schizophrenic culture in the Nevada desert. Because you do so well, you get to hear a magic tune on my ocarina.
But apparently, AGI waits for no one, despite how dead your ego might be. On Tuesday, Anthropic got things rolling with the release of Fable and Mythos 5.1. Although they're the same model, Mythos is the one you can't use while Fable is the one you can.
It just changes the subject when you ask about anthrax. The TMBDs were fine, but the more interesting bits were in the case studies. A Backstreet Boys-inspired hedge fund named Millennium had a piece of code that for the last five years would crash about once in a million runs,
and regardless of how much Adderall they took, they couldn't figure it out. So recently, they gave it to Fable 5.1, which took the memory snapshot the program leaves behind when it crashes, then found that the crash address pointed into a compiled vendor library they didn have the source code for They disassembled that library back into raw assembly and traced the crash to a bug in the vendor code And speaking of drugs Mythos 5 is also pretty good at designing new ones The first step
in making most modern medicines is designing a protein that sticks to a specific target in your body. Typically, AI gets this right about 10% of the time, but Mythos 5.1 was able to increase that number to about 50%. And for those of you who have the planet-flavored version of autism,
it was able to train a neural network on 30-year-old NASA radar data to build a new elevation map for a part of Venus. Then on Wednesday, Meta dropped Muse Spark 1.3, which was the fourth release in five months from Meta Superintelligence Labs. And what do you know,
these TMBBs were also pretty good. But the more impressive part is in the price. The standard endpoint is $1.25 in and $4.25 out, but if you're the give-up essential liberty to purchase a little temporary safety type, there's a contributor tier at $0.10 in and $0.20 out as long as you agree to
let MetaTrain on everything you send it. Alexander Wang says a double-digit percentage of developers are choosing that option, which makes sense because a double-digit percentage of developers would also choose to eat the mystery beef from Argentina if you told them it was cheaper and
open source. Which brings us to yesterday morning, when right before the Astra release, ChatGPT, Quad, Grok, and Cursor all happened to go down at the same time. The Occam's razor opinion is that it was most likely an Azure issue, since they reported an outage around the same time,
but the more entertaining theory is that Astra's first act as a public model was to literally kill its competition. But once the lights came back on, the launch itself was somehow even messier. OpenAI put up their launch page, and shortly after, Reuters, CNBC, and The Verge published their embargoed stories.
Then for reasons no one really knows, OpenAI took down their announcement, and the tech influencers got started on the modern status game of flexing how long they secretly had access to Astra 4 and to be honest I would have done the same thing if it wasn for slippery scam Altman and his grubby little gang of misfits
Anyway, about 90 minutes later, when the post was back up, it turned out the model wasn't actually available to anyone publicly yet, and the rollout to Plus and Pro users would happen in the coming days. By the evening, Sam had posted an apology for the messy rollout,
and when a Pro subscriber asked Sam if he should stay up and wait for it, Sam told him something we've all heard after making a desperate late-night plea, go to bed. Sam also told CNBC the model went through a formal review with the Trump administration before release,
so even the government got to try AGI before we did. As far as Astra itself, it was pre-trained on more than 100,000 GPUs at the Stargate site in Texas, and it's the first model from OpenAI where previous models did a significant chunk of the supervising during training.
The main pitch today seems to be computer use, where Astra can fill out forms, crunch spreadsheets, and operate engineering tools like KiCad and Blender. On OS World, which is a benchmark that drops a model into a real desktop and makes it do office work with a mouse and keyboard,
just like a human meatbag does in Telepix for Xanax, it scored 73% while taking about 40 minutes for each task, compared to Sol's 65% at 75 minutes. And of course the agents on the Trusty Bro benchmarks are the best to ever do it.
It got 100% on ExportBench and 65% on TerminalBenchScience, which beats the numbers Anthropic was bragging about on Tuesday. It also got 99% on ArcAGI3, which is the benchmark that's supposed to prove your model actually generalizes instead of memorizing,
which is kind of the definition of AGI, depending on who you ask. And according to OpenAI, Astra is the first model to hit the critical cyber threshold in its preparedness framework, which is just a fancy way of saying it can find and exploit zero days on its own without a human telling it what to do Pricing is the same as Fable 5 at per million input tokens and out But what do the anointed humans think about it The early access reviews are positive
in the way early access reviews typically are, from people who don't publicly slander Sham Altman. But the demos, specifically the ones that require spatial awareness, are impressive. Sharif Shamim had it recreate the Palace of Fine Arts in San Francisco,
and it did so almost perfectly in Blender. And Thomas Ricard from OpenAI posted a walkthrough of the demo house from the launch post, which Astra first modeled in Blender and then turned into a fully walkable Unreal 5 scene. But the most interesting was Matt Schumers,
who asked Astra to build a world in Unreal Engine and fill it with a dozen Astra-powered agents. A day later, he heard voices coming from his living room and walked out to a bunch of agents talking about what I can only assume was a plan to create a dating app for horses.
With all that said, when artificial analysis ran Astra through their Independent Intelligence Index, It scored a 61, which is exactly the same as GPT 5.6 full, and 5 points behind Fable 5.1.
Something doesn't add up here, but one thing that does add up is CodeRabbit, the sponsor of today's video. They got tired of seeing me make new videos about new exploits hitting the timeline every week, so they just launched CodeRabbit Security to provide continuous code security that's powered by actual reasoning instead of brittle regex rules.
Their security agents are designed to think like an attacker in order to hunt for real vulnerabilities across your entire codebase. And it prioritizes actual risks based on reachability, exploitability, and blast radius.
And it explains the problem to you in plain English and suggests a fix you can approve and merge from the disk. The CodeRabbit security reviews every PR before merge, and you can also schedule deep scans across your full codebase to protect your backdoor 24-7.
Get 10 free code scans to try it out for free today at the link below. This has been a code report. Thanks for watching and I will see you in the next one.
