AI's Big Week: Astra, Navier-Stokes & Security — Full Breakdown & Transcript

OpenAI talks GPT-6 Astra and Millenium Prize, researchers create WeWorm exploit & IBM’s US Open app

0h 33m video Published Sep 11, 2026 Transcribed Sep 15, 2026 IBM Technology IBM Technology
20.5K views Recent velocity 1.5 views/hour View full performance history →
Intermediate 8 min read For: AI enthusiasts, tech professionals, and enterprise decision-makers interested in AI applications and security.
AI Trust Score 70/100
⚠️ Average / Some Fluff

"The title promises a lot, and the episode delivers on all three topics, though the 'GPT-6' mention is misleading since Astra is not GPT-6."

AI Summary

In this episode of Mixture of Experts, the panel discusses OpenAI's new Astra model, its claimed solution to the Navier-Stokes millennium problem, IBM's AI-powered U.S. Open fan experiences, and a security exploit targeting WeChat. The conversation balances excitement about AI's capabilities with concerns about resource use, real-world applicability, and security implications.

[00:53]
OpenAI's Astra Model Announcement

OpenAI unveiled Astra, a new model with impressive demos including 2D-to-3D rendering and simultaneous actions like ordering Uber Eats. Panelists question the computational cost and real-world necessity.

[02:00]
Computational Concerns

Olivia Buzek worries about the token usage and whether such models are the right approach for agentic tasks, suggesting simpler solutions might suffice for many problems.

[03:37]
AGI Definition Debate

Aaron Blotman notes the lack of a converged definition of AGI and suggests it might appear as a 'really good digital employee.' He highlights the massive compute required to train Astra (100,000 Blackwell systems).

[05:49]
Customer-Centric Perspective

Brie Goffack states she hasn't seen a customer use case requiring such a large model, emphasizing that enterprise problems often don't need this scale, though speed is valuable.

[07:15]
Navier-Stokes Solved by AI

Astra solved the Navier-Stokes millennium problem using 10,000 AI agents, taking 88 hours and costing an estimated $15 million in compute, with verification via Lean in 17 hours.

[08:41]
Human-AI Collaboration

The solution involved human decisions, such as focusing on Navier-Stokes after an intermediate breakthrough. Brie emphasizes that human intellect and creativity remain crucial.

[14:51]
IBM's U.S. Open AI Features

IBM partnered with USTA to enhance the U.S. Open digital experience, using AI for match predictions, real-time likelihood-to-win updates, and a conversational Match Chat feature.

[21:20]
Limb Tracking and Biomechanics

IBM introduced serve quality tracking, monitoring 21 points on players' bodies at 50 fps, processing 1.2 billion data points across 254 matches to analyze biomechanics.

[23:28]
WeWorm Exploit Discovery

A security startup called Calif used AI to discover a memory corruption bug in WeChat's VOIP stack, leading to a zero-click worm exploit that could spread autonomously.

[27:36]
AI and Vulnerabilities

Olivia explains that all software has vulnerabilities, and AI makes finding them easier, but models still require prompts and are constrained by real-world resources.

The episode underscores AI's transformative potential in solving complex problems and enhancing experiences, but also highlights the need for careful consideration of resource costs, real-world applicability, and security risks.

Mentioned in this Video

💡 Key Takeaways

💡

AGI Definition

Provides a practical definition of AGI and questions whether we'll recognize it when it arrives.

03:53
🔧

Swarm of Agents

Illustrates a novel approach to problem-solving using a community of AI agents.

07:28
⚖️

Human-AI Collaboration

Emphasizes that human intellect and creativity remain essential in AI-driven discoveries.

12:17
📊

Biomechanics Tracking

Showcases a first-of-its-kind application of AI in sports analytics.

21:20
💡

Unconventional Attack Surfaces

Explains why AI exposes vulnerabilities in unexpected places, a key security insight.

27:36

[00:01] I haven't yet received a use case where my customers would have needed a model of this magnitude. All that and more on today's Mixture of Experts.

[00:17] I'm Tim Huang and welcome to Mixture of Experts. Each week, Emily brings together a panel of brilliant brains working at the frontiers of artificial intelligence to lead you through the week's news. On this week's episode, we've got Olivia Buzek, staff AI engineer, Aaron Blotman, IBM fellow, and Brie Goffack, the AI customer success engineer.

[00:36] And today, my co-host is David Zaks, staff writer for IBM Think. We're going to cover three big stories today. We're going to talk a little bit about IBM at the U.S. Open. We'll talk about some worrying news about WeWorm. But I think I want to start today by talking a little bit about all of this news we're seeing coming out of OpenAI.

[00:53] There's a new model, some major advancements in mathematics apparently. David, do you want to talk a little bit about this? Sure. So the model in question is Astra.

[01:06] It's the big new model from OpenAI. And if you've at a minimum seen the demo, these demos are becoming a staple of product announcements now. And often people go into a room and they begin talking with the model.

[01:19] and mainly I walk away thinking I really want a big kind of AI cave to kind of, you know, talk to a model in this way. But we've seen in this video incredible developments in terms of how they can turn 2D renderings into 3D renderings

[01:34] and simultaneously order Uber Eats and get that delivered apparently instantaneously. So just before we turn to the mathematical side of things, I would just be curious to hear from our theme panel on what they are excited about in Astra,

[01:47] whether this, as the NVIDIA CEO has suggested, is truly the AGI that we've been waiting for, whether it's a step change. What do you think, Olivia, for instance?

[02:00] So every time I look at things like this, like, you know, big new model coming out, I think a lot about the computational power that's going into it. So there is, yes, maybe it can do those things, but I'm wondering just how many essentially tokens are being used when it's doing something like controlling your computer?

[02:23] Like, is this really like the way that it makes sense for agent models to act in the world? And I think there's a lot of kind of low-hanging fruit that we have to do in terms of making agents actually accessible to the world.

[02:38] Sorry, making the world accessible to agents before we get overly excited about, well, it can solve an arbitrary problem. I can also order DoorDash. I could probably write a harness to convince a lower tier model to order DoorDash for me.

[02:55] In some ways, it's impressive. And in some ways, I just wonder if we're solving problems the right way. And so I think what worries me the most is like, are we approaching problems the right way as humans?

[03:10] Or are we simply throwing bigger and bigger models at problems without thinking through what the actual implications are? So it's a cool demonstration. I just wonder whether it has the right real world implications.

[03:25] Aaron or Bree, have you either tried the models or seen, tried Astra or seen a use case that you think is really novel or can be performed more effectively by this model than other models?

[03:37] Yeah, I mean, so first, I want to scratch that itch a bit about is this really AGI or not, right? Because, I mean, first of all, I think that we need to understand in the community, there's no converged definition of what AGI really and truly is, right?

[03:53] If we define AGI as a system that can autonomously take a novel goal, it can understand the problem, it can either go out and get missing knowledge, formulate a set plan, use tools, software, execute work, and so on,

[04:05] maybe that's AGI. But we may not even recognize when AGI has arrived. It just might look like a really good digital employee, just like one of us doing work. But I do hope we're at the beginning of AGI,

[04:18] because the reason is at what cost, right? I mean, if this number is right, It took 100,000 Blackwell NVLink 72 systems just to train this model. That's under order of enough power for an entire region in the continental United States.

[04:37] It's just unbelievable how many resources it took. Now, the other point I wanted to make is that the real AGI test shouldn't be a benchmark. That's known. It should be a system or goal that I haven't seen before to see what happens.

[04:51] And what we're all looking at here is the bench CAD benchmark, right? ASPR got an astounding 95.9% accuracy level on it. And the next highest one, of course, is the GPT 5.6 hole,

[05:06] which is what in the lower 80s and stable, roughly around the same place. So it did really good on that benchmark, right? And that benchmark is allowing the model. So I think circling back, David, to your question,

[05:20] And, you know, one task that is really interesting is this one, right, is giving the model multiple PD views, right, and having it generate executable CAD code and understanding that environment, building a 3D object, comparing it with real geometry.

[05:35] It's pretty neat that it can do those types of things. You know, and we're moving from, you know, being able to ask AI for answers to now asking AI to do work. Yeah, and let me add my perspective, which is very customer-centric.

[05:49] I haven't yet to see a use case, or as of today or as of this or last week, where my customers would have needed a model of this magic.

[06:01] We need to call the Navy or Stokes right now. And that's exactly the point. The things that we're solving on a day-to-day basis for, you know, for enterprises, it doesn't really require a model of that nature.

[06:16] I mean, could it make it somewhat faster? Yes, and speed is very critical in an enterprise setting. But I think I like what Olivia said, which we have to ask ourselves, what are we using this for?

[06:30] What problems are we really solving? But it's great that we have the capacity. So I think it's great we're pushing for those technological advances, be it computation, how many tokens are we burning?

[06:45] but then of course we have to ask ourselves what are we solving what can we what is the most efficient way to do it um but i have yet to uh throw this uh into a use case i'm sure it's

[06:58] probably going to be requested um and i'm going to play around with it at some point um but it's it's exciting to see what's happening i'm a huge nvidia fan i think they have they have a power and they're they keep pushing their own boundaries just like most other companies

[07:15] if you say. Yeah, one of the interesting aspects of the Navier-Stokes area isn't that AI solved the Navier-Stokes but it's, you know, and created the proof, right, but it's also that it used 10,000 AI agents

[07:28] to explore the problem, right, so it's this community of agents that work together with the coordinator almost like sub-agents that go out and respond and they can take an NCP server and get a tool, right, and run all these different explorations

[07:40] with chain of thought, right, but it's pretty cool, right, to see all of that work together, almost like a swarm of agents, right, that are now being unleashed on this. Maybe let's pause there just to,

[07:52] if you don't mind, Darren, to zoom out and just make sure viewers are up to speed on what we're talking about, this Navier-Stokes, Navier-Stokes, I've seen it pronounced any number of ways. But it is this so-called millennium problem in mathematics.

[08:06] It's in 2000, the Clay Institute, who said these seven problems in mathematics or like the new Mount Olympus in mathematics We should all focus on them And if you win if you solve them we will pay you a million dollars Only one had been solved in the past 25 years

[08:24] And the second one was solved. It's called Navier-Stokes. It has to do with fluid dynamics. I don't understand it. I call it the exploding water problem because apparently they've proven that water can spontaneously explode, which sounds awesome.

[08:41] but basically Astra solved it. There is some soap involved. There's kind of a soap opera because actually some humans at NYU slash Anthropic had maybe simultaneously solved it

[08:54] and it was only because of a rumor they were about to publish that OpenAI began to throw these 10,000 agents at the problem. But yeah, basically, Aaron,

[09:06] since you were kind of on this tack, do you want to tell us a little bit more about what the problem is how it was solved, the amount of, I gather it was sort of tens of millions of dollars worth of compute was thrown at this problem to win a $1 million prize.

[09:19] But what else should we know about Navy Reserves before we get into the weeds on sort of how it was solved and maybe if we have time, the soap opera? Yeah, I mean, like you said, this is one of the seven millennium prize problems.

[09:34] And it had been open for 90 years. So it's 90 years, almost a century this has not been solved. and what it's looking for it's looking at when a fluid velocity becomes unbounded

[09:46] even though the system starts very smoothly so like exploding water so it's very smooth and then you change the constraints and you figure out when does it become unbounded and the fluid velocity

[09:58] just changes and it's almost like it becomes sublime and it just changes right Is it safe for me to take a shower is what I need to know I would say it'd be interesting to take a shower.

[10:12] Safe? I don't know. It'd be a different kind of therapeutic shower. But now you have a proof that's been solved by Fable so that you can go look at it and ask the question, is it safe to take a shower?

[10:25] So it's pretty neat. But I do think we need to be careful. I mean, there is that soap opera, but even so, the mathematical community, it still needs to scrutinize the solution that came out

[10:37] right, of abstract here, right, because we need to make sure that it's universally accepted as the solution, right, and it's correct, right. So let's just wait and see what that ground truth actually is, you know.

[10:52] But I think the indicators, the litmus test, you know, it is looking pretty good, especially, you know, because other folks in the mathematics area are kind of up and on. So wait a minute, we solved it first, right. So it seems as though the signals are there, that it's pretty close.

[11:08] Brie, you have great expertise in terms of agent harnesses and how humans interact with agents. And I think a huge part of the Navier-Stokes story is about, it's not really that AI solved this problem,

[11:22] it's that humans managed AI to solve this problem. First of all, humans gave us the hunch of where to focus our energies. I gather that as Astra was turning on it, at a certain point it solved an intermediate problem.

[11:34] They were cranking on all seven millennium problems. At a certain point, Astra kind of solved an intermediate problem. Then the humans decided, let's throw all the energy at solving New Year's Stokes specifically.

[11:46] So a lot of human decisions in the loop as this was cranking. So Brie, as you read through this story, what were you thinking about this as sort of a collaborative human-AI effort? Yeah, I think that I love how you're framing this because it shows to the world that regardless of how powerful AI has become and is becoming, it still requires human intellect, human creativity, and human collaboration to get those answers to these huge problems.

[12:17] So I think that's a good thing. So for everyone who's watching and is scared you're going to lose your job next week because of AI, I don't think that's imminent. but we can have a different conversation about the future due to AI.

[12:32] But I think this is one thing we can learn from that. Our humans, our human brains are still incredibly relevant and also what we deem as important because the fact that this was thrown at Navier-Stokes

[12:47] and not something else, that tells us we as humans say this is important and we need to throw compute at it. The other thing that this elicited excitement within me is think about the advancements we're going to make in health and medicine.

[13:06] If we can use this, and I am convinced we will be solving, we will be finding cures to probably all cancers. I think I was listening to this podcast.

[13:19] I forgot the expert who was on. But he said, we will not be dying from diseases because AI and quantum physics and everything, that's going to help solve it. So, I mean, I think the initial solution took 88 hours.

[13:34] And then they verified, and I'm leaning on to what Aaron said, because I do believe they used another process where they translated the proof into lean, which is a, I believe, a mathematical.

[13:47] I'm not very familiar with the concept of weaning, but that only took 17 hours to verify whether the results of those 10,000 agents is in fact correct and whether we can start to believe it.

[14:02] I'm sure they're going to do additional verification around that, but I think this was incredibly exciting. this means we are a lot closer to cures and finding solutions to I mean think about global

[14:16] warming and you know things that were problems we're facing on a very major scale so those are my thoughts on this subject yeah it's a good note uh and an optimistic one which is good we don't always hear that about AI I'm going to move us on to our our next topic um long time listeners to

[14:36] Michelle will know that when Aaron appears, we usually talk about sports. And so we're going to move very briefly from curing cancer and global warming to what IBM has been doing at the U.S. Open. Aaron, I believe you want to talk a little bit about the demo that you've been working on, right?

[14:51] Yeah. So, you know, for more than 30 years. Right. So this is almost been since 1992. IBM and the USTA, we've partnered together to transform these digital experiences off the U.S. Open.

[15:03] So it's not bringing the tournament's data insights to the more than 14 million fans around the world. And just a fraction of them can actually attend. But this year, what we did is we took the partnership to another level by using AI.

[15:18] So some of these techniques that we are talking about, and I'll show you, what we did is we predicted matches before they even started. We would understand what's happening point by point, identifying the moments that changed the match as the match is being played.

[15:35] And then we also created a conversational and personalized experience for every fan so you could ask the burning questions that you had as tennis was unfolding. And then for the first time, which is really, really neat,

[15:47] it was one of those grand challenges that we took up, but we went all the way down and we traced and tracked limbs for athletes, biomechanics, and particularly on their serve.

[15:59] And so what I wanted to do was to just attempt to show you some of these demos that I have. So what you're looking at on my screen, if you go to, and it's live right now, the tournament is going on.

[16:13] It'll end this upcoming Sunday. But if you go to USOpen.org, and then if you go to the scores here, you can begin to see who's winning, who's doing what.

[16:26] And what I did is I picked the Pagoula match that just ended against Navarro And if I start at the top so let look at the preview So here our pre likelihood to win

[16:40] So what happens is we have the system that ranks all of the players. We have three max, you know, four casters that attempt to say what the volume of the media is being written about them,

[16:53] what the sentiment of different factors are. And then we also look at their head-to-head stats, if they have any. And then we also take into account different ratios, how are the players playing on each of the different surfaces,

[17:06] whether it's hard, clay, grass. But in essence, there's two classical machine learning models, so there's HD boost fees and logistic regression of the numbers.

[17:19] And they help predict who's supposed to win pre-match. And here, we got it right, where Bagula was supposed to win by 77%, and she did. It turns out that our odds were pretty good.

[17:32] I mean, we end up, I would say, in the 60s, with respect to predicting the right pre-match. And then, as I scroll down, I can see all the different stats that did happen. That just helps to give us the explainability.

[17:47] Now, if I want to look here, this is the replay. I can go in 3D, right, and then visualize and see how they're playing. But before I do that, I want to show you the summary of the match.

[17:59] So match chat is a place where I can go and ask questions, you know, during or after the match, so I can better understand what's happened. And then we have our life likelihood to win, right?

[18:11] So it's bootstrapped and started out with our pre-match likelihood to win. And then as the match goes on, you know, we have these momentum equations with different decays that change it.

[18:24] And then we also have different boosters based on the situation at hand. But it moves, you know, the probability that someone's going to win, you know, as every single point is made. So this is a real-time system.

[18:36] So as every single point is hit, won or lost, we get the message and then we compute, you know, the odds. and then at the same time after we do that we then send other messages over

[18:48] Confluent and we have a agentic marketplace that then takes the data and it measures whether or not this is a key moment and if it was then we want to

[19:00] produce text and here this one is the very end key moment but you can see the you know Pug Pagula closes out a brilliant two-to-one victory over Navarro But I can scroll back and just view all the key moments of the match so I can better understand it.

[19:16] And it's pretty cool. It's like a second screen experience. You can watch this as the match is actually going on. Now, going back to Match Chat, if I want to know what happened to the match, let's just push this button here, accept the terms.

[19:31] And then what's happening is this is exploring Match Chat. So it opens up another experience, and this comes and it queries through tools the data sources and feeds that we have, right,

[19:43] and we create a plot that has a context, right, and then it goes through, right, and it goes through this myriad of grass-laying type agents that three of them race, right.

[19:55] One of them is a pure Gen AI solution. Another has like a summarization, and then another uses a data synthesizer. synthesize it just because of the load scale that we have. If I want to see, you know, match details, then I can do that as well.

[20:09] Aaron, one question for you is I'm curious if you guys have information on what types of fans are using this type of experience. Because I feel like in my family I've got kind of like two kinds of fans represented.

[20:21] One of them just likes to tune in and watch the ball go back and forth. Yeah. I have another one who just really would love something like this, really is intense about the data, I'm just kind of curious about what you're learning about usage.

[20:33] Yeah, so we do, right? We have different user personas that we create these experiences for. What I showed you, like MatchChat, is for the middle of the road. It's not like a hardcore fan or the novice, but it's someone sort of in the middle, right?

[20:51] And maybe that fan wants to come or see the data to transition them into a hardcore fan. Or alternatively, maybe they're just too busy and they just want to see surface level and they can ask any questions that they want.

[21:03] But we do collect analytics and people do sign in and favorite players. And we're able to take that information and actually trace it back into types of questions in Match Chat to generate the types of personas that are using the system.

[21:20] So what you're looking at here, this is limb tracking. So we introduce serve quality, and we use IBM's Bob to build it. But what we do is we track 21 points across the player's body at 50 frames per second.

[21:34] So that turns out that over the course of the tournament, we're processing 1.2 billion data points and apply this to all 254 singles matches. And by doing so, we're capturing the biomechanics.

[21:47] So if I hit play, you can see the server hitting the ball. and we're measuring effectiveness, whether or not the ball goes into this polygon, and the efficiency or how they're swinging the racket. And we've broken up the serve into phases.

[22:00] So we have the start phase, a loading phase, clocking, whenever you get the racket up and you're about to hit the ball, acceleration when the racket's going down, to contact,

[22:13] to de-acceleration to finish and follow through. And we've broken out those phases and we apply rules to help us measure them. And, in fact, in the Match Chat experience, you can ask questions about the Asserve quality.

[22:26] And then on your mobile device, when you go to the homepage, we have this live experience that's personalized to you, and we show the Asserve quality and some insights around it. So it's pretty neat, and it's a first of a kind, and I'm really excited about expanding this to other types of shots.

[22:42] Nice. This is great, Aaron. Well, I love having you on the show because it feels like every time you come in, And the platform is advancing as we go, and we get to watch it together here on MRE, which is really cool.

[22:56] Well, Ray, I'm going to move us on to our sort of final topic of today. I think we've covered sports. We've covered curing cancer. And I think we're going to talk a little bit about the dark side for our final segment.

[23:09] A really interesting story, David, today that I think just broke over the last few days about the use of AI to generate viruses in effect, the thing called WeWorm that a lot of people have been worrying about. Yeah. So I mentioned now that AI has me worried about water and whether I'm safe to set foot in the shower.

[23:28] Now I'm really terrified of my phone because essentially this security startup called Calif used AI. to discover a bug in WeChat,

[23:42] which is the very popular Chinese messaging slash everything app. Over a billion people use it in China. And this Europe-based firm found this bug using their AI,

[23:59] and it's worse than a bug. What they were able to develop against this sort of vulnerability is a so-called worm, And not only that, but a zero-click exploit where essentially this virus can worm its way into one person's WeChat app autonomously, I believe, kind of place calls to other contacts.

[24:21] The recipients of that call don't need to do anything, and yet they are now infected, and it just can, in theory, grow exponentially. Now, this exploit was discovered in July.

[24:33] Of course, Akalop did the right thing. They went to the folks at Tencent, which makes WeChat. They alerted them to the vulnerability. The vulnerability was patched. Only now months later is the news story breaking But nevertheless I am personally I not a technical person but I terrified because this would seem like an exploit

[24:58] on the level of what we are more familiar with from, like, really sophisticated state actors. But it would seem that the rate at which AI is accelerating the discovery of bugs, the generation of exploits of these bugs,

[25:11] where should my P. Doom be at this point? I think you have a reason to be concerned. I think a private consumer of AI, you know, the end user, you know, we have, we interact with these tools.

[25:27] I personally have used WeChat during the two years I lived in Shanghai, so I'm very familiar with the app. I paid my rent using WeChat. That's how advanced it was back then in 2015-16 in China.

[25:42] So, and whenever I talked with our customers, security and implications of this, what we're discussing, is always part of the discussion.

[25:54] So, I mean, I wouldn't lose sleep over it, but it is definitely something we're seeing more and more every single month. I believe there's at least one major story about something like this and several smaller ones.

[26:07] I think the key takeaway here is that AI did make the hacking more powerful. I mean, a hacking case is always a significant event. It made it more accessible.

[26:19] So this is going to continue to happen, and I look to our security experts, which I am not one, to solve these issues around it. It's definitely a concern.

[26:31] A question I want to put to Olivia, if it's a little technical. I gather that we don't know exactly what the exploit was. Calif wrote that the bug is a memory corruption issue in WeChat's VOIP stack.

[26:49] And they also wrote something that this bug is an instance of the many unconventional attack surfaces that are present across many messaging apps. The other thing I noticed in their timeline is they said something like, our AI discovered the bug.

[27:03] And it's unclear to what extent they were just sort of letting the AI roam through various apps to find vulnerabilities. But it would seem to me that, and Bria sort of pointed to this, that the advent of AI, the AI era that we're living in, is exposing vulnerabilities in places we didn't quite suspect them, right?

[27:24] These so-called unconventional attack surfaces. Can you help explicate that a little bit to us here? Why are new corners of technology stacks suddenly discovered to be vulnerable in the AI era?

[27:36] Sure. So I've talked a little bit about this the last couple of times we've had big math problems getting solved by AI. So I think the way that I look at this, it's important to remember that we live in a dynamic system and an emergent system.

[27:54] There have always been rough edges in everything that humans have ever created. It's actually one of the core beauties of being human, is that the things that we create are not generally perfect objects.

[28:08] There's like entire lines of art, basically, about that fact, right? There's things like, you know, Wabi Sabi, where we put pottery back together and try to create a hole and things like that.

[28:22] There's a lot of imperfections. I know that sounds really philosophical, but what I want to bring that back to is all software does definitely have vulnerabilities of some kind.

[28:37] Some of them require more advanced things to take advantage of them, and others of them have very simple holes in them. But it's actually near impossible to build perfect software.

[28:50] So we can very much say that that's probably what's happening. But then I want us to take a look at these models. There has not been a case yet that I have seen where models are truly acting on their own.

[29:08] It is always in response to a prompt. And it's really important to remember that basically models do not have the capacity to act on their own. They are responding to text of some sort of prompt.

[29:25] Granted, as these models increase in their capacity, they're able to do more creative behavior in response to simpler and simpler prompts. So what may have been a while ago, hey, please explore the surface of this code base,

[29:44] look for this type of vulnerability and that type of vulnerability and this other type of vulnerability, may be as simple now as go off and look for vulnerabilities. But that lack of specificity that people are putting into their prompts now is coming back to them in the form of just the sheer amount of resources used.

[30:05] And so, yes, we can see things like this. But unlike with the Astra bit that we saw here, opening eyes, Astra, solving the Navier-Stokes problem, where they've actually reported the number of tokens that it required.

[30:26] And I think, as we've said, it took some $15 million worth of compute in order to solve a $1 million problem. Obviously, the economics of that just get weird at that point.

[30:38] But anyways, the point is it took a lot, right? We don't necessarily have that same data, or at least I haven't seen any, for this. What does it take to break into WeChat, right, to take advantage of these more complex ones?

[30:56] And my guess is the number is similarly quite high. So I think there's these real-world economics that there are, I think, essentially, and something I think we all have to kind of confront is OpenAI and Anthopic are behaving with so many resources available to them that they are able to apply these kinds of things relatively casually.

[31:17] And you can tell that it's casual, right? The way that they talk about it, it's like, oh, yeah, we just decided to go solve the millennium stokes problem. I don't have access to a $15 million budget. Like, I work in AI.

[31:29] I'm not allowed to go spend $15 million of tokens because I feel like the millennium power prize problem. Yeah, I couldn't do that. Like, even though I have access to these models, I can't go create worms, right?

[31:44] And so I think what we have to think about in this space is, yes, things have changed, right? We are in a new world, in new types of systems, but they are still very complex systems that are ultimately constrained by real-world resources.

[31:59] So personally, my fear is a little bit proportional to who holds those resources and what their priorities are. So anyone can do it as long as they have $15 million worth of compute to burn. Right, right.

[32:13] Which, like, yes, that's still a problem. But, like, in the era before, like, cybersecurity was still an issue. If you had $15 million to pay a hacker, you also could have done a lot of things.

[32:26] Okay. I'm willing to turn my phone back on now. Thank you very much. Wow. This is always a great panel to have on, Emily. We get a dose of hope, a dose of terror, a dose of grounded technical information.

[32:41] So, Brie, Aaron, Olivia, great to have you on the show. And David, it was great co-hosting you. And thanks for joining all your listeners. if you heard what you heard you can get us on apple podcast modifier and podcast platforms everywhere and we'll see you all next week on mixture of experts

More from IBM Technology

View all

⚡ Saved you 0h 33m reading this? Transcribe any YouTube video for free — no signup needed.