OpenAI Killed Its Most Loved Model
56sThe emotional attachment users had to GPT-4o and the shocking news of its deletion creates a strong hook.
βΆ Play Clip"The title promises a major shift, and the video delivers a detailed, data-backed analysis of that shift, though it's padded with a sponsor segment and some speculative content."
This video analyzes OpenAI's strategic shift away from the multimodal 'Omni' era, marked by the retirement of GPT-4o and Sora, and the launch of GPT-5.6 (Saul, Terra, Luna) focused on coding and enterprise work. It also covers Google's adoption of the 'Omni' brand with Gemini Omni Flash, Anthropic's focus on reasoning, and the competitive landscape of AI video generation.
GPT-4o was the first AI that felt alive to many users, serving as an emotional anchor. Its retirement matters because of this deep user connection.
OpenAI didn't pause or rebrand the Omni models; they deleted them piece by piece, while Google adopted the 'Omni' name for itself.
On January 29, 2026, OpenAI published a retirement plan. GPT-4o disappeared from ChatGPT on February 13, 2026.
Only 0.1% of users were still on 4o when it was pulled, but that equals about 800,000 people who lost their daily model in one afternoon.
The ChatGPT-4o latest API was cut off on February 16, 2026, forcing developers to migrate or watch their apps break.
On March 24, 2026, OpenAI announced Sora's shutdown. Disney found out less than an hour before the public. The app went dark on April 26, and its API dies on September 24.
OpenAI didn't just retire old models; they retired tools people had built entire businesses around, with zero migration path offered.
On April 3, 4o was wiped from custom GPTs, affecting business, enterprise, and education accounts alike.
In August 2025, OpenAI pulled 4o when GPT-5 launched, but users revolted and it was brought back within days.
22,000 people signed a petition to keep the old model alive, indicating a real dependency, not just nostalgia.
Voice mode still runs on 4o-based models under the hood. The API IDs GPT-4o Transcribed and GPT-4o Mini TTS are technically still alive, but the chat model and custom GPTs are confirmed dead.
The death of the chat model and custom GPTs is rated credible because it's based on OpenAI's own retirement notice, not a rumor.
On July 9, 2026, OpenAI launched GPT-5.6 with three models: Saul (flagship), Terra, and Luna. Saul tops the coding agent index with a score of 80, an independent benchmark.
OpenAI shipped programmatic tool calling, allowing models to chain tool calls without human confirmation, and ChatGPT Work, an enterprise agent for real workflows.
OpenAI is no longer targeting casual users; it's chasing enterprise seat licenses that scale revenue.
Altman claimed the new lineup is 54% more token efficient, but this is flagged as unverified because it's a company claim with no outside audit.
Saul, Terra, and Luna don't touch image generation, video generation, or voice synthesis. It's a purely analytical system for coding and enterprise work.
On May 19, 2026, Google announced Gemini Omni Flash at IO, three months after GPT-4o vanished from ChatGPT.
The video includes a sponsor segment for AI Master, a production pipeline for testing top LLMs, generating images, voice, and video in one place.
Google folded a version of Gemini Omni into Shorts, allowing anyone with a phone to generate video without touching the API. Free reach beats leaderboard rankings.
OpenAI coined 'omni' in May 2024 with GPT-4o, but two years later, they killed the brand and Google is wearing it now, signaling a changing of the guard.
Google confirmed a higher tier, Gemini Omni Pro, is coming but with no ship date. Rated as 'possible'.
The API runs at 10 cents per second of generated video, launched June 30. It's free inside Google's consumer apps, making it the most capable free video tool.
It accepts text, image, or existing clip inputs, outputs video, keeps character consistency, carries a synth ID watermark, and allows conversational editing.
Hassabis positioned it as a step toward a true world model, understanding video, audio, and images as a single continuous space.
Consumer video is a fast-growing category, and OpenAI left the building. Someone will fill that hole, and it won't be OpenAI wearing the Omni name.
Kwai Show's Kling 3.0 landed on February 4-5, with native 4K output and an Omni-branded variant. ByteDance's SeeDance 2.0 leads the benchmark.
Multiple companies are racing to own the word 'Omni' in the same year, signaling the category is still being defined.
Google kept V03.1 running separately, free in Google Vids. Runway's Gen 4.5 is for professional editors who want tight control over camera moves, lighting, and motion paths.
Sora died on April 26, and its API dies on September 24. OpenAI walked away from consumer video with zero replacement announced.
Some reporting suggests the old Sora team is pivoting to robotics research under a project nicknamed 'Spud', but OpenAI hasn't confirmed it.
Llama 3.14, LTX 2.3, Alibaba's 1.2.7, and MiniMax Helu 2.3 are also fighting in this space.
Opus 4.8 landed May 28, Fable 5 and Mythos 5 on June 9 (suspended 3 days later under US export control), and Sonnet 5 on June 30, now the default model.
Fable 5 scores 95.5% on SWE-Bench verified, beating Saul by ~15 points on SWE-Bench pro. On raw intelligence, Fable 5 is at 59.9 vs Saul's 58.9.
Saul wins on agentic coding tasks tied to tool calling; Fable 5 wins on raw reasoning and long context work. Pick based on product needs.
Claude dreaming previewed May 6, working like hippocampal memory consolidation. Harvey reported a ~6x lift in agent task completion, but it's a single customer report, rated 'possible'.
Anthropic chose to be the best reasoning model rather than compete for the Omni crown. No image, video, or voice synthesis in their lineup.
OpenAI is saving the number six for something bigger, a qualitative leap. The direction is long-term memory paired with autonomous agents, based on Altman's comments.
Prediction markets put a Q4 2026 launch as most likely, with a 28% chance it slips into 2027. Rated 'possible'.
Rumors of a 2 million token context window, code name 'Symphony', and 40% performance gains are unverified, from secondary blogs.
For casual use, subscriptions win on cost and convenience. Hardware pays off only for constant generation, privacy needs, or tinkering. The 2026 sweet spot is the RTX 5090 with 32GB VRAM.
When 2.7, Hunyuan Video, and LTX 2.3 are open-weight options. They won't match SeeDance or Kling on quality but run entirely on your own machine, offering privacy.
The ground keeps shifting. Build in a way that lets you swap models without rebuilding everything from scratch.
OpenAI invented 'Omni' in 2024, then spent this year killing it. Google now owns the word. Watch the brands, not the rumors, because the brands don't lie.
OpenAI has decisively ended the multimodal Omni era, pivoting to enterprise-focused analytical models, while Google has seized the 'Omni' brand with its consumer-friendly video tools. The AI landscape is shifting rapidly, and the key to navigating it is to build flexibly and watch the strategic moves of the major players, not just the rumors.
AI Master
tool
GPT-5.6 (Saul, Terra, Luna)
tool
Gemini Omni Flash
tool
Gemini Omni Pro
tool
Kling 3.0
tool
SeeDance 2.0
tool
Runway Gen 4.5
tool
Fable 5
tool
Mythos 5
tool
Opus 4.8
tool
Sonnet 5
tool
Claude Dreaming
tool
When 2.7
tool
Hunyuan Video
tool
LTX 2.3
tool
Llama 3.14
tool
Alibaba 1.2.7
tool
MiniMax Helu 2.3
tool
Sam Altman
person
Demis Hassabis
person
Harvey
service
What was the approximate number of users still on GPT-4o when it was pulled from ChatGPT?
About 800,000 people (0.1% of ChatGPT's user base).
01:00
What are the three models in OpenAI's GPT-5.6 lineup?
Saul, Terra, and Luna.
03:34
What is the name of Google's new multimodal model announced on May 19, 2026?
Gemini Omni Flash.
04:47
What is the API pricing for Gemini Omni Flash?
10 cents per second of generated video.
07:24
What is the name of the open-weight video model from ByteDance?
SeeDance 2.0.
09:06
What is the score of Fable 5 on SWE-Bench verified?
95.5%.
11:11
What is the name of the memory feature previewed by Anthropic on May 6?
Claude dreaming.
11:53
What is the most likely launch window for GPT-6 according to prediction markets?
Q4 2026.
13:30
What is the recommended GPU for local video generation in 2026?
RTX 5090 with 32 GB of VRAM.
14:39
The 0.1% User Statistic
It quantifies the real human impact of a seemingly small percentage, showing that 800,000 people lost their daily model.
01:00Saul Tops Coding Agent Index
It provides an independent, reproducible benchmark that validates OpenAI's enterprise focus.
03:34The Irony of the Omni Name
It highlights a major strategic shift, with Google adopting the brand OpenAI created and then abandoned.
06:56Fable 5 vs Saul Benchmarks
It offers a concrete, data-driven comparison that helps developers choose the right model for their needs.
11:11Watch the Brands, Not the Rumors
It's a practical principle for navigating the fast-moving AI landscape, emphasizing strategic moves over speculation.
15:34[00:04] just a model. For a lot of people, it was the first AI that actually felt alive to talk to. Some users called it their emotional anchor. That's exactly why what happened to it this year matters. OpenAI already answered the
[00:18] GPT-6 L question in a completely different way. They killed the entire Omni lineage this year. One product at a time. They didn't pause it or rebrand it, they deleted it piece by piece. And while they did that, Google walked in
[00:31] and took the word Omni for itself. I've got the full data timeline plus a credibility rating on every GPT-6 rumor out there. Stick around because by the end, you'll know exactly what ships next and what never will. OpenAI didn't try
[00:45] to do this quietly. On January 29th, 2026, they put the retirement plan in writing on their own blog. Two weeks later, on February 13th, GPT-4o disappeared from ChatGPT completely. And here's the number that should sting.
[01:00] Only 0.1% of users were still on 4o when it got pulled. That sounds tiny until you realize what it actually means. 0.1% of ChatGPT's user base is still around 800,000 people. They lost their daily
[01:16] model in one afternoon. Three days later, on February 16th, the ChatGPT-4o latest API got cut off too. Developers who built products on top of 4o had to migrate or watch their apps break. Then came the part nobody saw coming. On
[01:31] came the part nobody saw coming. On March 24th, OpenAI announced that Sora was shutting down. Disney found out less than an hour before the public did. By April 26th, the Sora app itself went fully dark. Its API dies for good on
[01:44] September 24th. The backlash inside the creator community was immediate. People who built entire workflows on top of Sora lost their generation pipeline overnight with zero migration path offered. That's the pattern running
[01:57] through this whole timeline. OpenAI didn't just retire old models, they retired the tools people had built entire businesses around. In between those two dates, on April 3rd, 4-0 got wiped from custom GPTs entirely. That
[02:10] hit business, enterprise, and education accounts alike. The Omni era wasn't fading out, OpenAI was surgically removing it department by department. And here's the twist most people missed. This wasn't the first time OpenAI tried
[02:24] This wasn't the first time OpenAI tried to kill 4-0. Back in August of 2025, they pulled 4-0 once already, right when GPT-5 launched. Users revolted so hard that OpenAI brought it back within days. When 4-0 came back around for good in
[02:39] February 2026, the backlash left a mark. 22,000 people signed a petition just to keep the old model alive. That's not nostalgia, and it never was. That's a real dependency, not sentimental attachment. So, when I tell you 4-0 is
[02:53] dead, I need to add a nuance most coverage skips. Voice mode still runs on 4-0 based models under the hood. The API IDs GPT-4-0 Transcribed and GPT-4-0 Mini
[03:05] TTs are technically still alive. Everything else you'd recognize is confirmed dead. That's the chat model and the custom GPTs. The consumer-facing Omni brand is gone, too. I'm rating that credible because it's OpenAI's own
[03:19] retirement notice, not a rumor. So, if OpenAI wasn't building GPT-6-0, what were they actually shipping? On July 9th, 2026, they launched GPT-5.6, split into three models: Saul, Terra, and Luna. Saul is the flagship and it
[03:34] currently tops the coding agent index at a score of 80. That's not a marketing number at all. That's an independent, reproducible benchmark, and I'm rating it credible. Alongside Saul, OpenAI
[03:46] shipped programmatic tool calling. It lets the model chain tool calls together without a human confirming every step. They also launched Chat GPT Work, an enterprise agent built for real workflows. It takes on an entire project
[04:01] working while you're away. That positioning tells you who OpenAI is casual users anymore. It's chasing enterprise seat licenses that actually scale revenue. Altman claimed the new lineup is 54% more token efficient than
[04:19] the previous generation. I'm flagging that one unverified because it's a company claim with no outside audit yet. Look at the pattern across this whole lineup. Nothing about Saul, Terra, or Luna touches image generation, video
[04:32] generation, or voice synthesis. This isn't an Omni model wearing a new name. It's a purely analytical system built for coding and enterprise work, and that's the real story of where OpenAI actually went. 3 months after GPT-4o
[04:47] vanished from ChatGPT, Google stood on stage at IO and announced Gemini Omni stage at IO and announced Gemini Omni Flash. That happened on May 19th, 2026. Google's reveal is coming up next. But first, here's the thing about this whole
[05:01] story. Every player I'm covering today lives behind its own login and its own pricing page. Each one needs a separate API key, too. To actually test GPT-5.6, Gemini Omni, and the rest for this video, I didn't open 10 separate tabs. I
[05:17] ran all of it inside AI Master. It's the same production pipeline my team and I use every day to [music] build our own channels. It's one window where I switch between any top LLM. I generate images, voice, or video in the same place, and
[05:32] every character stays consistent across all of it. We didn't build this as a separate product to sell you. We built it to run our own workflow, and now we're opening it up. There's a live community of over 12,000 paying users
[05:47] inside AI master right now. Every character you build there, you can monetize directly on the platform. Sharing is built in, too. You push any asset straight to your team or your audience without a second tool. The
[05:59] annual plan runs at a discount right now. If it's not for you, the 7-day money-back guarantee covers you with no arguments. Here's how you get access. Go to the link in the description and hit buy. Select the annual plan, fill in
[06:12] your details, and you'll get a confirmation email. Then log in and you're inside the same workspace I just showed you. The whole setup takes under 3 minutes. I ran the same prompt through Sol, Fable 5, and Grok 4.5 [music]
[06:28] back-to-back inside that dashboard. If you want to run side-by-side comparisons like this yourself, the link is below. All right, let's get back to Google's big reveal. Google folded a version of this directly into Shorts. Now, anyone
[06:43] with a phone can generate video without touching the API at all. That distribution move matters more than the benchmark score. Free reach beats leaderboard rankings when you're fighting for the next billion users. So,
[06:56] here's the irony nobody's really said out loud. OpenAI coined the word "omni" out loud. OpenAI coined the word "omni" back in May of 2024 with GPT-4o. Two years later, they killed the brand and Google is the one wearing it now. That's
[07:10] not a coincidence at all. That's a full changing of the guard. Google has also confirmed a higher tier is coming called Gemini Omni Pro, though there's no ship date attached yet. I'm rating that one as possible. Google said it's coming,
[07:24] but until it ships, treat it as a roadmap item, not a product. On the API side, it runs at 10 cents per second of generated video, which launched June 30th. But, here's the part that actually matters for most viewers watching this.
[07:38] It's free right now inside Google's consumer apps. That makes it the most capable free video tool anyone outside a lab has [music] ever had access to. Here's what makes Omni flash different from a normal video generator. You feed
[07:52] it almost any input, text, an image, or an existing clip, and it outputs video every time. It keeps character consistency across shots, and every output carries a synth ID watermark, so it's traceable as AI generated. You can
[08:07] then edit that video conversationally just by describing the change you want. Demis Hassabis positioned it as a step toward a true world model. His framing, one system that understands video, audio, and images as a single continuous
[08:21] space instead of separate bolted-on tools. That's the real crux of this whole story. Consumer video is one of the fastest-growing categories in AI, and the company that helped popularize it just left the building. Someone fills
[08:35] that hole, and soon it just won't be OpenAI wearing the Omni name when they do it. So, if OpenAI left the visual race, and Anthropic never entered it, who's actually fighting for video and image generation right now? Kwai Show's
[08:50] Kling 3.0 actually landed first on February 4th and 5th. It holds four separate entries in the artificial analysis top 10 with native 4K output and its own Omni-branded variant. ByteDance's SeeDance 2.0 came about a
[09:06] week later on February 12th and currently leads that same benchmark outright. Kling's Omni variant specifically targets the same any input-to-video workflow Google just shipped. Multiple companies are now
[09:19] racing to own that word in the same calendar year. That's rare, and it's a signal the category itself is still being defined. Google kept V03.1 running as a separate product from Omni flash. It's free inside Google Vids for anyone
[09:34] already in that ecosystem. Runway's Gen 4.5 remains the pick for professional editors who want the tightest control surface over every frame. Runway's control surface lets you pen camera moves, lighting, and motion paths
[09:49] instead of hoping the model guesses right. Professional editors pay for that precision because client work can't afford unpredictable output. That's a different customer than the one chasing free volume inside shorts, which brings
[10:01] us back to the hole in the middle of all this. Sora died April 26th and its API this. Sora died April 26th and its API dies for good on September 24th. OpenAI walked away from consumer video with zero replacement announced. One
[10:14] unverified rumor worth flagging here. Some reporting suggests the old Sora team is pivoting toward robotics research under a separate project also nicknamed Spud. That's a different Spud than the GPT-5.5 code name and OpenAI
[10:29] hasn't confirmed any of it publicly. I'm rating that one as unverified. It's also worth a quick mention that Llama 3.14 and LTX 2.3 are fighting in this same
[10:41] space. So are Alibaba's 1 2.7 and MiniMax Helu 2.3. Opus 4.8 landed May 28th and things moved fast after that. On June 9th came Fable 5 and Mythos 5.
[10:55] Both got suspended just 3 days later under US export control. Sonnet 5 followed on June 30th and is now the default model for free and pro users. Fable 5 is genuinely impressive on paper. It scores 95.5%
[11:11] on SWE-Bench verified and it beats Saul by roughly 15 points on SWE-Bench pro. On raw intelligence, the AA pro. On raw intelligence, the AA intelligence index has Fable 5 at 59.9
[11:25] against Saul's 58.9. If you're deciding which stack to build on for the back half of the year, this comparison actually matters. Saul wins on agentic coding tasks tied to tool calling. Fable 5 wins on raw reasoning and long context
[11:40] work. Pick based on what your product actually needs, not on which company shouts loudest. They did ship one memory feature worth mentioning. Claude dreaming previewed back on May 6th. It works like hippocampal memory
[11:53] consolidation, replaying past sessions to strengthen what the model retains. Harvey reported roughly a six times lift in agent task completion after adopting it. That's a single customer report, not an independent audit, so I'm rating it
[12:08] possible. That's not a gap in their strategy. It's the actual strategy they chose. Anthropic decided early that they'd rather be the best reasoning model on Earth than compete for the Omni crown. Every product decision since
[12:20] backs that up. Now, here's the part that actually matters for this video. Anthropic has no image generator and no video generator. There's no voice synthesis in the lineup, either. Claude design exists, but it's a workflow and
[12:34] document tool, not a Midjourney or Sora competitor. While OpenAI and Google fought over the Omni name, Anthropic shipped a completely different lineup and never once entered that race. So, let's finally answer the question
[12:46] everyone came here for. What is GPT-6 if it isn't an Omni model? Put it all together and OpenAI is saving the number six for something bigger, not just a version bump, an actual qualitative leap. That's why GPT-60 was never on the
[13:02] table. The O was already retired and OpenAI is reserving six for something it considers genuinely new. Based on comments from Altman himself, the direction is long-term memory paired with autonomous agents. These are
[13:16] systems that remember your context across sessions. They can carry out multi-step work without constant supervision. I'm rating that credible source, even without a firm date attached. On timing, prediction markets
[13:30] currently put a Q4 2026 launch as the most likely outcome. There's roughly a most likely outcome. There's roughly a 28% chance it slips into 2027 instead. That's a probability, not a confirmed date, so I'm rating it possible. Then
[13:44] there are more specific rumors floating around, too. Some people are talking about a 2 million token context window. Others point to a code name called Symphony or claims of 40% performance gains. None of that traces back to
[13:57] OpenAI. It's speculation from secondary blogs, so I'm rating it unverified until OpenAI actually confirms something. Quick detour here because people always ask me this in the comments. If subscriptions to all these tools feel
[14:10] expensive, can you just run video generation locally instead? Here's my honest verdict after testing both paths. For casual occasional generation, a subscription still wins on cost and convenience, full stop. Hardware only
[14:24] pays for itself in three cases. You're generating constantly, day after day. You need privacy over what you're creating, or you just enjoy tinkering with the setup. The 2026 sweet spot for that is the RTX 5090
[14:39] with 32 GB of VRAM. It runs somewhere between three and five thousand dollars depending on where you buy it. That's enough memory video models without constantly hitting out of memory errors. On the model side,
[14:54] out of memory errors. On the model side, When 2.7, Hunyuan Video, and LTX 2.3 are the open-weight options worth knowing about right now. They won't match Seedens or Kling on raw quality, but they run entirely on your own machine.
[15:08] There's also a privacy angle worth naming directly. Running models locally means your prompts and your footage never leave your own machine. That you wouldn't want sitting on someone else's server. If you're building on top
[15:21] of any of these platforms right now, the lesson isn't [music] which company wins forever. It's that the ground keeps shifting fast. Build in a way that lets you swap models without rebuilding everything from scratch. Open AI
[15:34] invented the word Omni back in 2024. Then they spend this year killing it piece by piece. The chat model, the custom GPTs, the entire consumer brand, and the company that owns that word now isn't Open AI. It's Google. And it took
[15:49] them exactly 2 years to take it. That single fact tells you more about where this race actually stands than any leaked screenshot ever will. Watch the brands, not the rumors, because the brands don't lie. If this saved you from
[16:02] chasing a model that's never shipping, subscribe. I'll keep tracking exactly who's winning this fight as it moves. Thanks for watching, and I'll see you in Thanks for watching, and I'll see you in the next one.
β‘ Saved you 0h 16m reading this? Transcribe any YouTube video for free β no signup needed.