[00:04] When my son Patrick was around three or four years old, and he was playing with these blocks with letters. I wanted him to learn to read eventually, [00:18] “Pa.” And he said, “Pa?” And then he said, “Pa-pa.” [00:33] And then something wondrous happened. "Pa! Patrick." [00:45] His eurekas were feeding my scientific eurekas. His doors, our doors were opening to expanded capabilities, expanded agency and joy. [01:00] Today I'm going to be using this symbol for human capabilities and the expanding threads from there which give us human joy. [01:14] Can you imagine a world without human joy? So I'm going to tell you also about AI capabilities and AI agency, [01:27] so that we can avoid a future where human joy is gone. I'm a computer scientist. My research has been foundational to the development of AI [01:41] My colleagues and I earned top prizes in our field, I'm not sure how I feel about that name, to talk to you about the potentially catastrophic risks of AI. [02:00] people have these responses. And I understand. How can this hurt us any more than this, right? [02:19] But recent scientific findings challenge those assumptions, To really understand where we might be going, About 15 or 20 years ago with my students, [02:34] and our systems were barely able to recognize handwritten characters. they were able to recognize objects in images. they were able to translate across all the major languages. [02:51] So I'm going to be using the symbol on the right in order to represent AI capabilities that had been growing In 2012, [03:04] tech companies understood the amazing commercial potential and many of my colleagues moved from university to industry. I wanted AI to be developed for good. [03:18] I worked on applications in medicine, for medical diagnostics, I had a dream. I'm with Clarence, my grandson, [03:34] and he's playing with the same old toys. And I'm playing with my new toy, the first version of ChatGPT. [03:47] because for the first time, we have AI that seems to master language. in every home. [04:00] this is happening faster than I anticipated, and I'm starting to think about what it could mean for the future. [04:12] but it might be just in a few years. And I saw how it could go wrong because we didn't, this technology eventually doesn't turn against us. [04:29] I'm a leading signatory of the "Pause" letter, where we and 30,000 other people asked the AI labs to wait six months [04:41] before building the next version. Then, with the same people and the leading executives of the AI labs, [04:53] And this statement goes: "Mitigating the risk of extinction from AI should be a global priority." [05:06] I travel the world to talk about it. and you'd think that people would heed my warnings. [05:18] I have the impression that people get this. Another day, another apocalyptic prediction. Hundreds of billions of dollars are being invested every year [05:36] And this is growing. of building machines that will be smarter than us, Yet we still don't know how to make sure they won't turn against us. [05:54] that the scientific knowledge that these systems have For example, by terrorists. the O1 system from OpenAI was evaluated [06:10] and the threat of this kind of risk went from low to medium, which is just the level below what is acceptable. But what I'm most worried about today is increasing agency of AI. [06:31] You have to understand that ... planning and agency is the main thing that separates us from current AI And these AIs are still weak in planning. [06:47] in this study, they measured the duration of tasks and it's getting better exponentially fast. What are AIs going to do with that planning ability in the future? [07:05] Recent studies in the last few months have tendencies for deception, and maybe the worst, self-preservation behavior. [07:22] So I'm going to share with you a study that is helping us understand this. the AI has read in its input And we can see in its chain of thought [07:36] that it's planning to replace the new version After it executes the command on the computer, And the AI is now thinking how it could answer [07:53] And it's trying to find a way to look dumb, for example. And it's a lie, a blatant lie. [08:05] What is it going to be in a few years There's already studies showing that they can learn in these chain of thoughts that we can monitor. [08:23] they would not just copy themselves on one other computer They would copy themselves over hundreds or thousands of computers But if they really want to make sure we would never shut them down, [08:38] they would have an incentive to get rid of us. So I know I'm asking you to make a giant leap into a future that looks so different from where we are now. [08:52] To understand why we're going there, there's huge commercial pressure to build AIs with greater and greater agency to replace human labor. [09:05] We still don't have the scientific answers, You'd think with all of the scientific evidence we'd have regulation to mitigate those risks. [09:20] But actually, a sandwich has more regulation than AI. So we are on a trajectory to build machines that are smarter and smarter. And one day, it's very plausible that they will be smarter than us, [09:36] and then they will have their own agency. which may not be aligned with ours. [09:48] Poof! We are blindly driving into a fog, that this trajectory could lead to loss of control. [10:06] Beside me in the car are my children, my grandson, my loved ones. Who is beside you in the car? [10:19] Who is in your care for the future? The good news is there is still a bit of time. We still have agency. [10:34] We can bring light into the haze. My team and I are working on a technical solution. It's modeled after a selfless, ideal scientist [10:51] without agency. that are trained to imitate us or please us, which gives rise to these untrustworthy agentic behaviors. [11:04] we might need agentic AIs in the future. So how could the Scientist AI, which is not agentic, fit the bill? The Scientist AI could be used as a guardrail against the bad actions [11:22] of an untrusted AI agent. that an action could be dangerous, You just need to make good, trustworthy predictions. [11:36] the Scientist AI, by nature of how it's designed, for the betterment of humanity. We need a lot more of these scientific projects [11:50] and we need to do it quickly. Most of the discussions you hear about AI risks [12:02] Today, with you, I'm betting on love. Love of our children can drive us to do remarkable things. [12:17] (Laughter) I'd rather be in my lab with my collaborators, working on these scientific challenges. [12:29] and to make sure that everyone understands these risks. We can all get engaged to steer our societies in a safe pathway [12:44] in which the joys and endeavors of our children will be protected. I have a vision of advanced AI in the future governed safely towards human flourishing for the benefit of all. [13:04] (Applause) Thank you. (Applause and cheers) [13:23] one question. In the general conversation out there, a lot of the sort of fear that people spoke of is the arrival of AGI, [13:38] What I hear from your talk The right thing to be worried about is agentic AI, But hasn't the ship already sailed? [13:54] There are agents being released this year, almost as we speak. it would take about five years to reach human level. but we still have a bit of time. [14:09] The other thing is, we have to do our best, right? We have to try because all of this is not deterministic. If we can shift the probabilities towards a greater safety for our future, [14:24] CA: Your key message to the people running the platforms right now is slow down on giving AIs agency. to understand how we can get these AI agents to behave safely. [14:42] And all of the scientific evidence in the last few months point to that. YB: Thank you.