TubeSum

Gemini 3.5 Pro Delay — Full Breakdown & Transcript

Google's Gemini 3.5 Pro Delay: What Really Happened Inside DeepMind

0h 16m video Published Jul 27, 2026 Transcribed Aug 14, 2026 A AI Master
Intermediate 8 min read For: Tech enthusiasts, AI industry professionals, and investors following Google's AI strategy.
AI Trust Score 70/100
⚠️ Average / Some Fluff

"Delivers a well-researched breakdown of the Gemini delay with source badges, but the sponsor segment and some filler reduce density."

AI Summary

Google's flagship Gemini 3.5 Pro model remains undelivered months after its promised June release, wiping $425 billion in market value amid rumors of internal struggles. Despite a blowout quarter with $119.8 billion in revenue, the company has missed two public launch windows, leading to widespread speculation and analysis of what went wrong inside DeepMind.

[00:00]
Gemini 3.5 Pro Delay

Google promised Gemini 3.5 Pro for June, but as of late July, no ship date has been announced. The delay has erased roughly $425 billion in market value on rumors alone.

[00:28]
Bloomberg Reporting on Delay

Bloomberg published a report on July 16th based on 10 current and former Google employees. The core claim is that Gemini 3.5 Pro underperformed on internal coding benchmarks, leading to a retrain in late June that returned disappointing data.

[01:36]
Gemini 3.6 Flash Released

On July 21st, Google shipped Gemini 3.6 Flash, a cheaper and faster model but not the flagship. It scores the same 50 on the Artificial Analysis Intelligence Index as 3.5 Flash, indicating no intelligence gain.

[06:05]
DeepMind Researcher Exodus

In a six-day window in June, DeepMind lost five senior researchers, including Nelm Shazier (original Transformer author) and Nobel laureate John Jumper. This exodus is credible and publicly documented.

[09:31]
Leaked Specs Analysis

Rumors about 3.5 Pro include a 2 million token context window (possible), DeepThink mode (credible but unverified for 3.5 Pro), pricing at $15/$60 per million tokens (unverified), and deep agent/computer use integration (possible). Google has published zero ethical benchmark numbers for 3.5 Pro.

[12:06]
Google's Strategic Thesis

The video argues Google is deliberately ceding the raw intelligence crown to win territory in distribution and cost. The Gemini app has 950 million MAU, the API processes 22 billion tokens per minute, and 90% of the Fortune 100 uses Gemini Enterprise.

[12:56]
Apple Deal Details

Apple pays Google about $1 billion per year for a custom 1.2 trillion parameter Gemini model to power the new Siri, slated for iOS 27 in September 2026. This will reach 1.4 billion iPhones.

[14:35]
Gemini 4 in Pre-training

Sundar Pichai confirmed on the Q2 earnings call that Gemini 4 is already in pre-training, calling it 'very ambitious' and promising a near-monthly release cadence. No date for 3.5 Pro or Gemini 4 was given.

Mentioned in this Video

Study Flashcards (12)

What was the promised release month for Gemini 3.5 Pro?

easy Click to reveal answer

June

How much market value did Google lose on rumors of the delay?

medium Click to reveal answer

Roughly $425 billion

What was the core reason for the delay according to Bloomberg's reporting?

medium Click to reveal answer

Gemini 3.5 Pro underperformed on internal coding benchmarks.

02:31

How many current and former Google employees did Bloomberg speak to for their report?

hard Click to reveal answer

10

02:31

What model did Google ship on July 21st instead of 3.5 Pro?

easy Click to reveal answer

Gemini 3.6 Flash

01:36

What is the pricing for Gemini 3.6 Flash per million input tokens?

medium Click to reveal answer

$1.50

07:50

How many senior researchers did DeepMind lose in a six-day window in June?

medium Click to reveal answer

Five

06:05

Which company did Nobel laureate John Jumper leave DeepMind for?

hard Click to reveal answer

Anthropic

06:19

What is the score of Gemini 3.6 Flash on the Artificial Analysis Intelligence Index?

medium Click to reveal answer

50

08:48

How much does Apple pay Google per year for the custom Gemini model powering Siri?

hard Click to reveal answer

About $1 billion

12:56

What is the target release for the Gemini-powered Siri?

hard Click to reveal answer

iOS 27 in September 2026

13:26

What did Sundar Pichai confirm about Gemini 4 on the Q2 earnings call?

medium Click to reveal answer

It is already in pre-training.

14:35

💡 Key Takeaways

💡

Bloomberg's Core Claim on Delay

Provides the most credible, sourced explanation for why Gemini 3.5 Pro is late: poor coding benchmark performance.

02:31
📊

DeepMind Researcher Exodus

Documents a rare, concentrated loss of top talent including a Nobel laureate, signaling internal turmoil.

06:05
📊

Gemini 3.6 Flash Intelligence Score

Reveals that Google's latest release is not smarter than its predecessor, only cheaper and faster.

08:48
💡

Google's Strategic Thesis

Presents a counterintuitive argument that Google is deliberately ceding the intelligence race to dominate distribution and cost.

12:06
📊

Apple Deal Details

Highlights a massive, underreported distribution advantage: 1.4 billion iPhones powered by Gemini.

12:56

[00:00] Google promised Gemini 3.5 Pro for June. It is now the end of July and the model still has no ship date. Over the last several weeks, Google wiped roughly $425 billion in market value on rumors alone.

[00:15] And here is the part that makes no sense on paper. Google just posted a blowout quarter with $119.8 billion in revenue. The flagship went missing while the business printed money.

[00:28] Every AI channel on YouTube is guessing at what happened inside DeepMind. I actually read the Bloomberg reporting behind the headlines. So in this video, every leak gets a badge on screen.

[00:40] Credible means named reporting or an official Google statement. Possible means plausible that's been resourced. Unverified means aggregator noise, and I will call it out by name. By the end, you will know exactly what is real and what Google is actually betting on instead.

[00:56] Let's rebuild the timeline with actual dates, because the dates alone tell most of this story. On May 19th, Google took the I.O. stage and shipped Gemini 3.5 Flash on the spot.

[01:08] The pro model got a promise instead, and that promise said next month. So the entire industry circled June on the calendar. June came and went with silence from the Gemini team.

[01:20] There was no blog post and no revised ship date. Then reporting pointed to an internal target of July 17th. That day passed too, and Google said nothing publicly. On July 21st, Google finally shipped something that it was not the model anyone was waiting for.

[01:36] Gemini 3.6 Flash arrived and said, and we will break down exactly what it is in a few minutes. Which brings us to today, July 24th. There is still no ship date for Gemini 3.5 Pro,

[01:49] and Google has quietly stopped naming targets. That is a flagship missing its public window twice in a row. For a company that later promised a near-monthly release cadence,

[02:01] this silence is the loudest data point we have. One more detail from ILO matters later in this video. Google's own session materials describe serious pro capabilities, including a 2 million token context window.

[02:15] Keep that in mind, because those materials are the strongest source behind several of the leaks we are about to race. Here is the strongest piece of reporting we have on why the model is late. On July 16th, Bloomberg published a report by Julia Luff and Beatty Alba.

[02:31] They spoke to 10 current and former Google employees, which makes this the best source account anywhere. The core claim is about coding. According to those sources, Gemini 3.5 Pro underperformed on internal coding benchmarks.

[02:46] Coding is exactly the market Google can least afford to lose right now. The same reporting described the retrain and run in late June that returned disappointing data. So the delay is not mysterious, and it is not a conspiracy.

[03:01] The model just was not clearing the bar Google set for it. I rate all of that credible, because it comes from many reporters with ten sources. Now let's talk about what the aggregator channels did with that story.

[03:13] You have probably seen videos claiming the base model was scrapped and retrained from scratch. That claim is unverified, and nobody has traced it to an actual source. You have also seen the phrase third delay in a row.

[03:26] That framing is also unverified, because only two missed targets are actually documented. Some channels claim hallucination-driven failures killed the launch. There is no source reporting behind that one either, so it gets the same unverified batch.

[03:42] And I am not the only one calling this out. The tech outlet Neumburn flagged both the scrapped model claim and the hallucination claim as unsourced. When a claim only exists inside YouTube thumbnails, that already tells you what it is worth.

[03:57] Every claim in this video got checked against primary sources. Not just Bloomberg, Google's own materials and the earnings called Transcript 2. I didn do that by manually digging through browser tabs I ran the research and the fact check through my platform The model testing happened right there too inside AI Master Here what that actually looks like

[04:18] Every major model lives in one window. Gemini 3.6 Flash, GPT 5.6 Sol, Cloud Fable 5, all of them side by side. I pulled up the same prompt about Gemini's coding benchmark.

[04:31] Then I ran it through three models at once, right here in this dashboard. That's not a gimmick for this video. That's the actual workflow behind every script on this channel. But the model picker is just one piece.

[04:44] And I don't want you walking away thinking that's all this is. Every top LLM lives here, and generation runs cheaper per token than going direct. Image, voice, and video generation left in the same window.

[04:56] Same pricing advantage applies there too. You build a character once, and it stays consistent across every generation after that. You can also publish that character and monetize it directly on the platform.

[05:08] Everything you generate, you can share straight out. There's also an academy inside the platform. So you're not just getting the tools, you're getting shown how to actually use them and practice immediately. It works for your reach instead of just sitting in a folder.

[05:23] There's a live community of over 12,000 paying users inside AI Master right now. This isn't a tool you use alone. It's people trading feedback and building an audience together. None of that feedback is scripted. It's pulled straight from our own demo reel.

[05:39] The annual plan runs at a discount right now. If it's not for you, the 7-day money-back guarantee covers you with no arguments. Here's how you get access. Go to the link in the description and hit buy.

[05:51] Select the annual plan, fill in your details, and you'll get a confirmation email. Then log in, and you're inside the same workspace I just showed you. The whole setup takes under 3 minutes. All right, let's get back to Google's story.

[06:05] While the model flipped, something worse was happening to the team behind it. Over roughly a six-day window in June, DeepMind lost five senior researchers. Nelm Shazier, one of the original Transformer authors, left for OpenAI.

[06:19] Then Anthropic picked up four more in quick succession. John Jumper went first, and Bloomberg separately reported Jonas Adler and Alexander Kritzel moving together on June 24th. Arthur Cagney followed a day later.

[06:32] Jumper is a Nobel laureate for Alpha Fold, so this was not junior attrition. This exodus is credible and publicly documented, and every one of those names is confirmed. For context on how unusual this is, look at Signal Fire's talent data.

[06:47] It says deep mind researchers are nearly 11 times more likely to leave for anthropic specifically than the reverse. Right now, deep mind sits on the wrong side of that ratio.

[06:59] The market reaction tells you how much narrative matters here. On June 22nd, the Exodus News wiped roughly $225 billion off Alphabet's market cap. Then the Bloomberg Delay Report on July 16th erased another $200 billion or so.

[07:15] That is the $425 billion number from the top of this video, and it moved on stories rather than earnings. One thing I want to keep clean, because a lot of coverage gets it wrong. The roughly 5% after-hours dip on July 22 came from CapEx guidance of $195 to $205 billion.

[07:36] It had nothing to do with the model delay. So what did Google actually ship on July 21? Gemini 3.6 Flash is real, and every number here is credible because it comes from Google's own launch materials.

[07:50] Pricing lands at $1.50 per million input tokens and $7.50 per million output tokens. Compare that to 3.5 Flash, which charged $9 on output alone. You still get the full 1 million token context window at that price.

[08:05] The model also runs leaner, producing roughly 17% fewer output tokens on the same tasks. Fewer tokens at a lower rate means your real bill drops twice. The benchmark gains are real too at least in specific lanes On DeepSWE the agenda coding Benchmark it jumped from 37 to 49 It hits 83 on OS World Verified which measures computer use

[08:31] On MLE Bench, it scores 63.9% on machine learning engineering tasks. Those are meaningful jumps for a point release, and I tested the coding lane myself. For high-volume agent work, this model is genuinely a better deal than anything Google had before.

[08:48] But here is the sentence that reframes the entire launch. On the Artificial Analysis Intelligence Index, Gemini 3.6 Flash scores a 50. That is the exact same score as 3.5 Flash, the model it replaces.

[09:03] Google made the model cheaper and faster rather than smarter. That is a completely legitimate engineering win, but it is not a flagship. Google also refreshed the low end with 3.5 slash like and a security-focused 3.5 slash cyber variant.

[09:18] Both of those follow the same pattern, and they optimize cost instead of raising the ceiling. Now for the part you clicked for, the leaked specs. I am going rumor by rumor, and every single one gets a badge.

[09:31] Remember the ground rule from the research on this. Google has published zero ethical benchmark numbers for 3.5 Pro. So anyone showing you a precise 3.5 Pro benchmark chart is making it up.

[09:45] Rumor 1 is the 2 million token context window. This one actually traces back to Google's own I.O. session materials, which described it directly. But describing a capability on stage is not the same as shipping it.

[09:59] No launch configuration has been confirmed anywhere. I rate this possible, and honestly it is the most likely spec on this risk. Rumor 2 is DeepThink mode. DeepThink itself is credible because Google shipped an update to it back in February 2026.

[10:16] It exists, and people use it every day. That part is not in question. The specific claims about how DeepThink performs inside 3.5 Pro are a different story. Every number attached to that pairing is unverified, so the batch splits down the middle.

[10:31] Rumor 3 is pricing at $15 input and $60 output per million tokens. Nobody has traced that figure to the source inside Google. It is a plausible price point given the market, but plausible is not source.

[10:46] This one stays unverified until Google publishes the pricing page. Rumor 4 is the one I think everyone underrates. Multiple signals point to deep agent and computer use integration in the pro model.

[10:59] Look at what Google already shipped as evidence of direction. That 83% OS World Verified score in 3.6 Flash is a computer use benchmark. If Trell extends that range, Google differentiates on something and frottage and open AI have not locked up.

[11:16] I rate the integration possible and it is the rumor I would actually watch. Let's put the whole board on screen for one honest minute. On the Artificial Analysis Intelligence Index as of late July,

[11:28] Claude Fable 5 leaps at around 60. GDT 5.6 Sol sits right behind it at roughly 59. QEK 3 holds third place at about 57,

[11:40] which still surprises people. In Google's best available model, Gemini 3.6 Flash scores the 50. Read that leaderboard again and notice what is missing. Google has no model in the top tier right now.

[11:52] Full stop. Its would-be contender is the one without a ship day. I am keeping this comparison short on purpose, because the interesting story is not the lease this month. The interesting story is why Google seems weirdly calm about losing it.

[12:06] Here is my actual thesis, and it is the reason this video exists. I do not think Google is losing this race by accident. I think Google is deliberately ceding the raw intelligence crown

[12:19] to win territory nobody else is defending. The crown gets you headlines, but the territory gets you revenue. Start with distribution because the numbers are absurd The Gemini app sits at 950 million monthly active users The API is processing around 22 billion tokens every minute Roughly 90 of the Fortune 100 runs on Gemini Enterprise

[12:42] an AI mode and search across a billion monthly users. And Floppy would trade a lot of benchmark points for any one of those lines. Then there is the Apple deal, which I think is the most underpriced story in AI right now.

[12:56] from our German Bloomberg reporting from November 2025. Apple pays Google about a billion dollars a year. In exchange, a custom 1.2 trillion parameter Gemini model powers the new Siri

[13:08] running on Apple's own private cloud compute. And let me kill a myth on screen right here. You may have heard this already launched in spring 2026, and that claim is simply false. The public target is iOS 27 in September 2026, and it is still labeled a beta.

[13:26] When it does land, Gemini instantly reaches 1.4 billion iPhones. That is distribution no benchmark score can buy. The second pillar is cost, and this is where the TPU story matters.

[13:39] Google's new TPU 8i targets roughly 80% better inference price performance than Ironwood, and Google is not just using these chips, it is selling them to rivals. Anthropic committed to about a million Ironwood TPUs over five years,

[13:54] a deal reportedly worth up to $40 billion. Meta is a TPU customer too. Half the industry is quietly hedging against NVIDIA, and Google owns the hedge. The third pillar is multimodal generation.

[14:09] Geo 3.1 and Omni are genuinely competitive in their categories, even against fast-moving Chinese rivals like Kling and Seedance. Google isn't dominating video generation outright,

[14:21] but it doesn't need to. It just needs multimodal to be good enough to keep users inside the ecosystem. The other two pillars do the heavier lifting anyway. So what actually comes next? And what did Google commit to on the record?

[14:35] On the Q2 earnings call, Sundar Pichai confirmed that Gemini 4 is already in pre-training. He called it very ambitious and he promised a near monthly release cadence going forward. That confirmation is credible because it came straight from the CEO on the recorded call.

[14:52] what it does not include is a date for Gemini 4 or for 3.5 Pro. September is the next beat that genuinely matters. That is the iOS 27 window

[15:04] when the Gemini-powered Siri is slated to reach iPhones and data. Reporting suggests a possible wait list at launch, so temper your expectations on day one. But strategically, it is the moment Google's distribution thesis

[15:18] either pays off or slips again. Meanwhile, the rest of the industry has started openly taunting Google. After the Flash-only drop on July 21st, Metas Alexander Wang tweeted two words,

[15:31] Gemini who? Axios reported on July 23rd that rival executives were quick to show the releases. Here is the uncomfortable framing I keep coming back to. Third parties are naming Google's game right now, and Google is letting them.

[15:45] Its answer is not a single benchmark reader. It is the whole compute and distribution stack. Whether that answer works is the story we will be tracking through September. Let me land this with three things you can actually act on.

[15:59] First, stop waiting for 3.5 Pro to fix your workflow. Use 3.6 Flash for cheap high-volume work because the economics are genuinely great. When you need frontier reasoning, use Cloud Table 5 or GBT 5.6 Solve today.

[16:16] Watch for a firm Gemini 4 date against that monthly cadence promise. A real day is the recovery signal, and vague ambition is not. Third, if you are on iPhone, September is your beat. Siri running on Gemini changes distribution overnight, and distribution is the whole thesis.

[16:33] If this breakdown saved you from another aggregator rabbit hole, subscribe for the next one. Every leak on this channel gets a badge, not a guess. I'm Arthur, and I'll see you in the next one.

More from AI Master

View all

⚡ Saved you 0h 16m reading this? Transcribe any YouTube video for free — no signup needed.