Million-Token Context Windows: A Gimmick?
44sProvocative claim that no one uses million-token context windows sparks debate among AI enthusiasts and developers.
▶ Play Clip"The title asks a question the video answers directly, but the content is thin and conversational, lacking depth."
The video discusses the practical use of million-token context windows in large language models, questioning why developers rarely utilize them despite their availability. It highlights issues like context rot, cost, and the mismatch between context window size and enterprise-scale data needs.
Despite million-token context windows being available for two years since Gemini introduced them, most users stay under 200k tokens.
The more context fed to an LLM, the lower the quality of output generally becomes, a phenomenon known as context rot.
Users are charged for every context token, so voluntarily increasing context length makes the model more expensive.
Even a perfect million-token context window is insufficient for enterprise data, which can start at 8 trillion tokens, making 100x improvements irrelevant.
The video concludes that million-token context windows are largely unused due to context rot, cost, and the vast scale of enterprise data, making them impractical for many real-world applications.
How long have million-token context windows been available?
Two years, since Gemini first introduced them.
00:01
What is context rot?
The phenomenon where the more context you feed an LLM, the lower the quality of output becomes.
00:15
Why does using a larger context window increase cost?
Because you are charged for every context token, so more context means higher cost.
00:29
What is the typical token size of enterprise document databases?
They can start at 8 trillion tokens.
00:55
Million-token context windows are rarely used
Challenges the hype around large context windows by pointing out actual usage patterns.
00:01Context rot degrades output quality
Identifies a key technical limitation that affects practical LLM usage.
00:15Enterprise data scale dwarfs context windows
Highlights the mismatch between context window sizes and real-world data volumes.
00:55[00:01] lengths of LLMs a lot further. Like and and my I obviously I might regret saying this, but we've had million token context windows for 2 years now since Gemini first introduced it, and no one uses it. Everyone stays under 200k. Is
[00:15] of context rot and the fact that the more context you feed it generally the lower the quality output will become? So So yes, that that is that is up there uh also cost reasons, right? Like you get charged for every single context token
[00:29] you voluntarily make it more expensive for yourself? Uh and then there's also just order of magnitude reasons, right? Like let's say context window today with zero context rot. It's perfect recall across all of
[00:41] it, right? Absolutely solved. 1 million. Then I Then I come along and go, you million token context window. Uh so you just 100x and like also perfect context you know, uh amazing. Great. But I'm sitting over here. I'm in enterprise,
[00:55] and my document uh database and storage and notes started at 8 trillion tokens. What is your 100x going to do for me? Nothing. [laughter] Absolutely nothing.
⚡ Saved you 0h 01m reading this? Transcribe any YouTube video for free — no signup needed.