The AI That Could Be Our Last Invention
42sThe historical tie-in and bold claim of an ultra-intelligent machine being our last invention immediately hook viewers with a thought-provoking premise.
▶ Play Clip"The title is a bit sensational, but the content does deliver on the core promise of explaining RSI and the new 'dreaming' technique, though it spends a significant portion on a sponsor segment."
This video from The Code Report covers a new paper from ByteDance, Tsinghua, and other Chinese labs called 'The Last AI Built by Humans,' which outlines a five-stage roadmap for Recursive Self-Improvement (RSI). The video breaks down the core concept of 'dreaming'—where an AI agent tests thousands of policies against cached runs of its own past attempts—and compares it to a similar paper from Google DeepMind and the University of Maryland. It also evaluates whether this process truly constitutes RSI or is just a more efficient search algorithm.
I.J. Good's 1965 prediction that the first ultra-intelligent machine would be the last invention man needed to make, as it could improve itself recursively. This is called Recursive Self-Improvement (RSI).
33 researchers from ByteDance, Tsinghua, and other Chinese labs published a paper laying out a five-stage roadmap for RSI, where the final stage involves the AI rewriting its own improvement process.
Google DeepMind and the University of Maryland responded with a paper called 'Wet Dream RSI', claiming that by turning an AI's old discovery logs into a simulator, the AI could 'dream' up thousands of new search strategies and improve without touching the model itself.
The key insight is that the exploration policy—the decision of what to try next—was previously hard-coded. The new approach saves every attempt (code, score, crash logs) so a new policy can be tested against old runs without needing to re-run the model.
The 'dreaming' process allowed the agent to test thousands of policies against the same run and keep the best one. In tests, it wrote a lasso solver that beat Python's standard ML library in ~300 tries, compared to 550 for a static policy and 51,000 for the previous record holder.
The prompt is crucial: it instructs the agent to read past attempts, avoid tiny tweaks, and 'pinky promise' not to kill processes. This highlights the importance of prompt engineering in guiding the agent's behavior.
By I.J. Good's definition, no. The model writing the new policies is still the same Gemini, so it can't find solutions it wasn't already capable of writing. It just finds them faster. This is true of all recent AI math breakthroughs—they use static models wrapped in a custom harness.
The video concludes that this is 'just a cool search algorithm with some caching,' not true RSI. The sponsor, Blacksmith, is then promoted as a drop-in replacement for GitHub Runners that is faster and cheaper.
The video argues that while the 'dreaming' technique is a significant efficiency gain for AI problem-solving, it does not constitute true Recursive Self-Improvement. The model itself remains static; only the search strategy improves, making it a powerful but ultimately limited approach.
The RSI Dream
Explains the foundational concept of RSI and its historical context, setting the stage for the entire video.
Dreaming as a New Technique
Introduces the core innovation of the video: using cached runs to test new policies without retraining the model.
00:42Impressive Results
Provides concrete numbers showing the efficiency gains of the 'dreaming' approach, making the claims more credible.
02:49The RSI Verdict
Provides a critical analysis of whether the technique is truly RSI, offering a balanced perspective.
03:16[00:00] In 1965, a straight British mathematician named I.J. Goode, who had spent World War II fighting Nazis next to a gay Alan Turing in Bletchley Park, wrote that the first ultra-intelligent machine would be the last invention man ever needed to make because once an AI gets good
[00:14] enough to improve itself, every improvement makes it better at improving, which makes it better at improving, which you get the idea. This is called RSI, and it's been the wet dream of AI researchers ever since. Well, last week, 33 researchers from ByteDance, Tsinghua, and a few other Chinese labs
[00:29] published a paper called The Last AI Built by Humans. It laid out a five-stage roadmap for RSI, where in the final stage, the AI rewrites the process it used to improve itself, and we all get reassigned to the countryside for agricultural work.
[00:42] Then on Sunday, Google, DeepMind, and the University of Maryland responded by dropping a similar paper of their own called Wet Dream RSI. In it, they claim that by turning an AI's old discovery logs into a simulator and letting it dream up thousands of new search strategies inside of it,
[00:56] the AI got better at discovering things without anyone ever touching the model itself. In today's video, we'll break down how Dream RSI works under the hood and decide whether an agent being able to rewrite its own exploration policy is actually RSI or just more hype slop It is September 17th 2026 and you watching The Code Report Every time AI has made a mathematics breakthrough the process has been the same one that Alpha Evolved popularized
[01:20] last year. You take a coding agent, hand it a problem and a scoring function, then run a loop where it proposes a solution, evaluates it, reads the feedback, and tries again a few thousand times. This process is what discovered the Jacobian conjecture this summer, and what OpenAI used to
[01:35] front-run the Navier-Stokes problem earlier this month. But the most interesting part of the process is one no one really talks about called exploration policy. The idea is that at every step of the loop, there's a decision to make about what the agent tries next. For example, say the loop is optimizing
[01:50] for a matching algorithm for horses. If one attempt pairs up horses slightly better than others, does the agent keep building on it, or does it start over with something that could be better? And if an attempt crashes before a single horse gets matched, is it the whole idea that's bad,
[02:03] or just the implementation? Until this week, this exploration policy was hard-coded into the loop by whoever set it up. But what the DeepMind team figured out is that if you save everything from every attempt, like the code it wrote, the score it got, and whether or not it crashed, you don't
[02:16] actually need to touch the model again to test a new policy. You can just show the new policy the old runs that are cached on the disk and let it decide where to go from there And because that costs nothing the agent can test thousands of different policies against the same run and then keep whichever one would have reached the best result in the fewest attempts Then it deploys that policy on the next run saves that run too and repeats the process again
[02:37] The paper calls this process dreaming, and to test it, they pointed Gemini at eight different problems across algorithm design and mathematics, then ran the same setup with a fixed policy to see if the dreaming version could beat it.
[02:49] I won't bore you with the TMBBs, but the most impressive one was that it wrote a lasso solver that beats Python's standard machine learning library in about 300 tries, where the static policy needed 550 and the previous record holder needed about 51,000.
[03:03] And because we're living in hell, the most interesting part was the prompt. It basically begs the agent to read every past attempt before writing any code, to stop making tiny tweaks to the same idea over and over, and to pinky promise not to kill any processes.
[03:16] So is any of this actually RSI? By I.J. Good's definition, no, because the whole point is that the thing doing the improving gets smarter each round. In this case, the model that writes each new exploration policy is still the
[03:28] same Gemini, so it can never find a solution that it wasn't already capable of writing. It just finds them faster and with fewer wasted attempts. But with that said, that's also true of every AI math breakthrough we had this year The Jacobian conjecture Navier and progress on the Ryman hypothesis all came from static models wrapped in a custom harness with sub swarms and orchestration doing most of the heavy lifting
[03:50] and the weights only get better when a human goes back and trains the next model on what the swarm found. So if you gave IJ mushrooms and convinced him that a human in the loop still counts, he might agree that this is the last invention man ever needed to make.
[04:02] But when he came down, he'd probably argue that this is just a cool search algorithm with some caching. And that's why you need to check out Blacksmith, the sponsor of today's video. It's a drop-in replacement for GitHub Runners that lets you run your GitHub actions twice as fast while costing 75% less.
[04:16] And they just launched Codesmith, a cloud coding agent that knows your repos and CI runs, so you can ask it to build something from GitHub, the web, or Slack like I'm doing here. I'm asking it to add a new AI provider to my app and wire up the API key in my infrastructure repo,
[04:32] and it can open a PR in each one, then fix failing tests or address review comments without relying on messages between bots. You can also ask it to recommend better runner sizes for your CI history and turn those changes into a pull request, so you're not renting a supercomputer to check your
[04:48] semicolons. Try it out for free and get 3,000 GitHub Actions Minutes at the link below. This has been The Code Report, thanks for watching, and I will see you in the next one.
⚡ Saved you 0h 04m reading this? Transcribe any YouTube video for free — no signup needed.