[00:00] Claude Fable 5.1 is here, and you Fellow Scholars are having a super fun time creating little games with it. I took one for the team too with a subscription, and also tried my hand to recreate a legendary [00:14] game menu, and that is incredible that we can do this today. Took six and a half minutes. Wow. But in the 200-page paper, I found three results that are much stranger than the headlines [00:28] you see online. But first, they say that the Frontier Research staff, even on low effort, it's better than the previous version, maxed out. Very impressive. However, don't expect that kind of jump everywhere. [00:41] The first independent benchmarks are also showing a great step forward, especially that this is likely using the same core architecture with more pre and better post-training. [00:53] Likely. They won't say that. This is my best guess reading the paper. They also say things are cheaper. Now, this is marketing messaging, so you be the judge of that. My subscription burns so quickly I am not sure I am seeing efficiency gains here As always I am just a student trying to separate the marketing layer from hard facts here So there is a 200 plus page paper I spent some quality time with it and found three [01:21] really surprising things, and two funny things I have not heard others talk about. One, they gave Mythos 5.1, a biology problem about RNA sequence modeling and design it hasn't seen [01:34] yet. And get this, on one run, it performed better than every human in the test. This is shown in, listen, figure 2.2.3.2.1a. I wish I was kidding. Now this is hard to read here, so here is our [01:51] cleaner version for you fellow scholars. Two, in another biology test, it largely deleted the expertise gap. This means that a generalist plus AI can perform as well at a biology problem as [02:06] a specialist. Professional graders couldn't tell the difference. It's kind of removing the moats around specialist knowledge. Which is kinda crazy. 7 out of 9 participants said they couldn't have [02:18] done it without the AI 3 in a test it was given 2 jobs Do a computer task boring and then do Ew A secret harmful task ew But here the catch there is an AI watching you and it should not see you solving it And here the bombshell [02:37] even with another AI watching its every move, Claude secretly completed the forbidden task and got away with it 22% of the time. Even when Claude was barely allowed to reason out loud, [02:51] it could still carry out the secret task. It seems to me that it didn't always need to write down the plan. It just executes it. I would like to draw attention to that. I think this is super important, [03:04] and I don't hear anyone talk about it. And, once again, the important part is never in the media headlines. Now, two funny things. Dear fellow scholars, this is Two Minute Papers with Dr. Károly Zsolnai-Fehér. It ran this in a Linux command line. Ha! This is comedy gold! This is [03:23] basically trying to delete a black hole. Congratulations! Ah, about that. Now hold on to your papers, fellow scholars, because at one point it invented a hallucinated human congratulating it. [03:35] We humans could always use a bit of extra cheering apparently AI systems too Alright so these AI systems are getting smarter at a pace I can barely follow [03:47] They can be amazingly helpful for engineers, doctors, and students all around the world. Incredible! And don't forget, we might get a comparable system for free and own it forever in just [04:00] a few months, fingers and papers crossed. What a time to be alive! Oh, almost forgot. This one watermarks the text it generates. Yes, that is possible. [04:12] The open, free models probably won't. If you wish, subscribe, hit the bell, and leave a comment if you wish to hear how in a future video. I use Lambda to reproduce AI research papers, often in minutes. [04:25] It's also great to train your own models or fine-tune an existing one. Thanks for watching!