GPT-6 Astra Harder To Monitor, AI Model Mania, and Gemini’s Mountain Fail
YouTube will ask you to confirm your subscription.
About this episode
GPT-6 Astra is putting up some eye-watering benchmark scores, with several well into the 90s. Justin’s impressed enough to consider ditching Anthropic. Frank’s less dazzled by the scores and more worried that OpenAI’s most capable model is also getting harder to monitor.
Plus: Nvidia’s $12 billion Unicode joke, the US government weighing in on OpenAI’s copyright case, Gemini 3.8 Flash becoming old news within days, Meta’s absurdly cheap Muse Spark 1.3 coding model, Fal generating video faster than you can watch it, Tencent shrinking HY4 without making it stupid, and Gemini sending three hikers up Mount Shasta rather badly prepared.
JOIN THE ARGUMENT
Are smarter AI models becoming harder to monitor?
Sources
- There's an Easter egg hiding inside the $12.9303 billion price tag of the Nvidia-Hugging Face deal
- US government sides with OpenAI on issue of training LLMs on copyrighted material
- Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- Claude: Fable 5.1 and Mythos 5.1
- Introducing Muse Spark 1.3
- Introducing H3 Max by fal
- Tencent released open-source Hy4 preview model
- Introducing GPT‑6 Astra
- Researchers fear safety disaster ahead of OpenAI’s Astra release
- Inside OpenAI’s Reboot
- Hikers rescued after using AI to plan Mount Shasta hike
Key insights
Are smarter AI models becoming harder to keep an eye on?
As AI gets more capable, its internal reasoning may also become harder to monitor. That creates an awkward trade-off: the systems we most want to trust could be the ones giving us less visibility into how they reached a decision. Better performance may be arriving alongside weaker monitorability.
Would reaching AGI actually change the world overnight?
Even if a model crosses the line into AGI, businesses still have to rebuild processes around it, people still have to learn how to use it, and scientific breakthroughs still take time to reach the real world. The technical milestone could arrive long before the economic transformation does.
What if AGI is still stupid in very weird ways?
A system could outperform people at most economically valuable work and still confidently send you up a mountain with the wrong plan. That contradiction may not disappear with better models. The real challenge could be learning when extraordinary AI capability still needs ordinary human judgement.
TranscriptThis transcript was generated with AI and may contain errors.Read full transcript
Frank: I’m your resident doomer, Frank Prendergast. With me is, as always, techno-optimist Justin Collery. And Justin…
Did Nvidia really make a $12 billion joke?
Justin: Oh, to be rich, Frank. Oh, to be rich. Quick joke for you. Oh, to be rich, right? So we talked about—last week we talked about Hugging Face.
Frank: Yeah.
Justin: Bought by Nvidia for, we said at the time, 12 billion quid, thereabouts. Here’s the nerdy joke, you’re gonna love this one, right? It wasn’t bought for exactly 12 billion quid.
Hugging Face is called Hugging Face—do you know why it’s called Hugging Face?
Frank: I actually do not.
Justin: All right, so there’s an emoji when you do the—
Frank: Oh yeah, the—
Justin: It’s a little face and it goes like that, right? And it’s a—
Frank: Oh, that one. Yes.
Justin: Yeah, yeah, yeah, yeah. That’s why it’s called Hugging Face. It’s named after that emoji. It was sold for $12 billion, 930 billion, $300,000.
So one, two, nine, three, oh, 300, right? Turns out that one, two, nine, three, oh, 300 is the Unicode character for the hugging—
Frank: No. No.
Is this true?
Justin: Yeah, the price was a joke. They changed the price to make it the Unicode character for the hugging face.
Frank: The price is a joke.
Justin: Isn’t that incredible? Oh, to have so much money that you can pay for stuff in Unicode.
Frank: Just—and make a joke out of, like, $12 billion. Say, oh my God, that’s crazy. That’s insane.
Justin: Anyway, somebody else is very rich.
Frank: So I… What? Where…?
Justin: On.
Was Frank right about OpenAI and copyright?
Frank: Well, I just wanted to tell you that the producers were onto me, and they’ve got a new segment for the show that they want us to put in this.
Justin: Oh, excellent.
Frank: Yes, and the name of the segment is Frank Was Right.
Justin: Oh, okay. This is a short segment.
Frank: It is. It is actually a short segment. So did you come across the story where the lawsuit that The New York Times filed against OpenAI, saying that, “You can’t, you shouldn’t have trained on our data, it’s copyright material,” et cetera, et cetera? So that’s still ongoing. We still don’t have a resolution on that.
But I saw a story saying that the government, the Trump administration, had contributed a 20-page brief which basically says, “Please, courts, you must allow OpenAI to train on copyright material. We need a competitive advantage. America needs the competitive advantage of AI, and this is a matter of national security.
We don’t want China getting ahead,” et cetera, et cetera, et cetera.
Justin: Wait, just before we get to the bit where you say that Frank was right, okay? Hold on a second. So the American government has filed a thing to say that, “Your Judge, don’t be bold here. Do the right thing by OpenAI.” Anthropic are also getting sued this week by Sony Music and I think it’s Warner Brothers for huge mass, industrial-scale taking away of our intellectual property.
Do we think that the American government will also give a 20-page brief to that judge to say, “Do not—or do the right thing by Anthropic?”
Frank: Yes. We have blacklisted Anthropic and we’ve deemed them a supply chain risk, illegally, but all the same please do allow them to use copyrighted material. Yes, I absolutely see that coming, yes.
Justin: All right. Okay. We’ll see what happens there. Anyway, carry on. Where were you? Right.
Frank: So I—well, does this sound familiar to you? This, oh, it’s a matter of national security to be allowed to train on copyright material because way back when, last year, March 2025, I think it was episode 48 of The AI Argument, OpenAI released this essay that they were giving to the government, saying everything that should happen with AI.
In there, there was something that I took exception to, which was you must allow AI companies to train on copyright material because it’s a matter of national security ’cause we don’t want China getting ahead. And I said at the time, this… I hate this kind of manipulation and political manoeuvring where they’re clearly trying to get the government to intervene, and I did not think it would take this long.
I didn’t think we’d see the black-and-white evidence over a year later, but here we are now and the government have basically repeated the same thing in a letter to the courts.
Justin: But does that not mean that you’re wrong? You’re not right, you’re wrong.
Frank: No, as in—
Justin: What that means is that it is a matter of national security importance ’cause the government got involved in a court case, meaning that the letter that OpenAI wrote was totally correct. It is a matter of—they wrote a letter saying, “This is a matter of national importance,” and the government has now agreed with them by signing this letter, therefore making them right and you wrong.
Frank: What’s the—like, isn’t this what we talk about all the time in terms of regulatory capture? Isn’t this it at the highest level you could almost imagine? I’d go back to what I said at the time of, look, if it’s a matter of national security, if it’s at that level, do you know what?
We gotta take this out of the hands of the private companies, and you know what? Let’s just nationalise them.
Justin: Interestingly, that is in the Overton window these days, right? So there are growing calls to nationalise some of the companies. Anyway, so that’s… Well done.
Frank: Yes. Thank you. Thank you. Yeah. Oh, yeah. I should’ve lined up a sound effect of, “Woo, yay, Frank was right. Woo.”
Justin: There we go. Anyway…
Are AI models now coming like buses?
Frank: For everything.
Justin: So we’ve been saying for the last four weeks, we are in the singularity. Welcome to the singularity, people, day 31. And this week, normally I would have some sort of medical thing or science thing or a maths thing or whatever it is. There’d be all sorts of things going on.
This week models have been like buses. We didn’t get one model release, we didn’t get two model releases, we didn’t get three model releases, we got about 10 model releases. All of them were groundbreaking in their own special way, except for one or two which were totally boring.
Did Gemini’s new model become outdated in days?
Justin: Which ones do you want to talk about first? The totally boring ones or the ones that were groundbreaking and interesting?
Frank: Let’s get the boring ones out of the way. Let’s even see if I agree with you that these are the boring ones.
Justin: All right, so the first boring release of this week from Google, we had the release of Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. Boring.
Frank: Fascinating. I’ve literally—the reporting on this that I’ve been seeing is like, “Oh, Google are back.” People wrote them off and now they’re back. Excellent model, not frontier-level model but, you know, I guess a little bit like we were talking about… I can never remember this model name, GLM Flash 5.3.
A little bit like we were talking about that. It’s remarkably capable for the size of model it is, and it’s remarkably cheap. So last week we talked about GLM, and that’s a Chinese model, and we were saying that even though it’s impressive and it’s cheap, people may not wanna use a Chinese model.
So I think the Gemini 3.8 Flash is interesting from that perspective. It’s quite capable. It’s very, very cheap, and if you don’t want to use a Chinese model, this might be for you.
Justin: Everything that you just said there was correct on Monday. It was relatively cheap and relatively close to the frontier. That was true on Monday. By today, unfortunately, it is no longer relatively cheap and it is no longer relatively close to the frontier because of all the other models that were released Tuesday, Wednesday, Thursday, Friday.
So…
Frank: The speed, the speed at which things change.
Justin: And I’m not even joking, right? So the second boring release this week comes from my favourite AI company, Anthropic. They released Claude Fable 5.1 and Claude Mythos 5.1, which nobody can touch anyway. And it’s like, yeah, whatever, they’re both just too expensive for most people to use, and they talk too much, and they speak a language that it’s really hard to understand.
It’s like, yeah, great, whatever, not that interesting.
Frank: I mean, that just blows my mind that Fable 5.1 and Mythos 5.1 could be released and we’d be like, “Yeah, boring.” But in this case, I tend to agree with you. Allegedly they’re better at long-running autonomous research, and that’s brilliant and it’s great, but it’s like, yeah, in terms of what I need to do, how I want to use AI, it’s like, yeah, no, not even relevant to me.
Justin: All right, great.
Is Meta now the cheapest way to code with AI?
Justin: So let’s move on to the more interesting releases which we’ve had this week, in no particular order at all. So the first one is Meta released Muse Spark 1.3. What is Muse Spark 1.3 and why is it interesting? It is a coding model. It is primarily used for coding. It is released by Meta, your favourite social media company, and it’s brilliant, right?
It is frontier quality. It is as good as Opus 5, or there and thereabouts, right?
Frank: Oh wow.
Justin: A little bit better than Gemini 3.8 Flash that you were speaking of earlier. The big selling point for this, if like me you’ve got loads of kids, is its price. So the input and output cost on this model is five cents per million tokens, so long as you’re a contributor.
I think it’s five cents input and I think 15 cents output per million. It’s basically free. It’s basically free, and it’s apparently brilliant, right? And I will spend my weekend using it, and if it’s—and it comes, by the way, with their own CLI, which you can use to write code, so its own version of Claude Code, its own version of Codex, and I’m gonna give it a go.
And if it’s as good as everything I’ve read, that will be the daily driver for all of my children. It’s like, just go use that one because it is so cheap you could not get to the end of your weekly allowance dollar, whatever. And they also do a plan for $20 where you can—they give you an allowance, which I assume also you could just not get to the end of ’cause you get so many tokens.
Frank: And exactly as you said there, I mean, at that price, again, if you don’t wanna be using a Chinese model, this sounds like, well, this would be—this is an American model that would be… I mean, it’s cheaper than either the Gemini one or the GLM one. That’s—
Justin: Incredible. Brilliant. Big news. Good—go, go Mark Zuckerberg, good for you. Things you didn’t expect to hear on a Friday evening.
Can AI now make video faster than you watch?
Justin: Now, the next big release this week, we should have a new section, you know, a companion section to your moments when Frank was right: moments when Justin was wrong.
Many, many years ago, I made a prediction that we would have on-the-fly video generation within 12 months. I did a sort of a straight-line calculation. I was wrong, Frank. I’ll admit that I was wrong. It’s taken us two years, not one year, to get to this point, but Fal have done it. So a company called Fal has released a video generation.
The videos are amazing. They’re realistic. They’re photorealistic quality. And there’s loads of… I think the Jerry Seinfeld memes that you can see on the internet nowadays are generated using this model. But the main thing about this model is it generates video faster than you can watch it.
So you can sit down, you can say, “I want to watch this,” and it will start to stream the video to you, and it will just keep going. In theory, you could watch infinite video.
Frank: That is—
Justin: Stream the video to you faster than it can make it. That is, to me, an inflection point in AI video. It changes—people are doing calculations at the moment, right?
If you were to stream constantly, right? So if you were to do The Truman Show, the AI—
Frank: Yeah. Okay. Yeah.
Justin: Constantly, right? As long as you can make more than $30,000 a month off it, you’re making money.
Frank: That’s mental. That is mental. Speaking of which, do you happen to know what the costs of the generations are?
Justin: Well, you can figure it out. If it’s 30,000 a month, right? That’s about 1,000 a day, 24 hours, so divide that 1,000 by 20 is 100. It’s about, what’s that? 250 quid? I don’t know.
Frank: I’m just, I’ve—the meme, I’m—that meme with all the numbers and the person doing puzzles, that’s me right now.
Justin: Sometime we can figure it out afterwards. It’s not cheap, it’s not expensive, right? It’s only gonna get cheaper.
Frank: That is crazy. That is crazy. Yeah, and that is definitely a model I’m gonna have some fun exploring.
How did Tencent shrink AI without dumbing it down?
Justin: All right, so this was a little bit more esoteric, right? It’s a little bit more—before we get onto the other big one this week. So a company called Tencent have released a model, HY4, and I just want to describe to you what it does. I think this is super clever. So they’ve released a model. The model started off being 1.5 terabytes, which is big for the non-techies.
It’s a big model. They squeezed it down to 200 gigabits, so they made it, whatever that is, like a factor of six, eight times smaller. How did they do it? It didn’t lose any intelligence. They did a thing which is so blindingly obvious that I’m surprised nobody has done it before. So normally what they do for, let’s say, the Gemini 3.8 Flash, these flash models, and, you know, Haiku, and you can download—
They do this thing where they—they call it—they have a word for it, right? But basically, the numbers that represent all the weights in these models, they’re normally, let’s say, an accuracy of eight decimal places, so there’s eight numbers after the decimal point. And what you can do is, if you wanna make it smaller, you just reduce that accuracy to maybe six numbers or maybe four numbers, right?
And they do that across the entire model, and that for sure halves the size of the model, but it also makes it a little bit stupider because those weights decide what words you get on the output. What these people at Tencent did was super smart. So what they did was they got a whole corpus of questions, and they got this huge model, 1.5 terabytes, and they put all the questions through it and watched which neurons in the AI lattice or network activated, and they kept all those ones and just took out all the other ones or made all the other ones a lower precision.
And the result is you get a model which is way smaller, but when you ask it the important questions, it’s just as smart as the big model, which seems like an obvious solution and really cool. So—
Frank: Wow.
Justin: Go Tencent, right? That was just a techy nerdy one.
Has GPT-6 Astra broken the benchmarks?
Justin: All right, onto the big release, the one we all want to talk about.
Frank: Yes. So my favourite company and one of my least favourite CEOs, OpenAI.
Justin: Yes. They have released GPT-6, otherwise known as Astra, and the benchmarks are off the charts. I mean, we’re talking off the charts. This—now, we can’t use it yet, right? We can’t get access, not just because we’re in Europe, but we’re also not on their whatever favoured supplier programme, whatever it happens to be.
Frank: This is another one of those—so this is kind of, this is OpenAI’s Mythos moment in a way. It’s very, very cyber capable, and so it’s only being rolled out right now to people on their special cyber development platform programme that they have. And then they’re saying they will roll it out to the rest of us eventually.
So yeah, I think it mirrors what happened with Mythos and Fable, doesn’t it?
Justin: Yeah, well, except that they seem to be promising that they’re gonna roll it out. I mean, for me—they’re indicating that everybody is going to get access to this model, whereas we still don’t have access to Mythos unless you’re on a special programme, so there is that. I mean, look, I just gotta jump into it, right?
There’s all sorts of… This model, the numbers are just eye-watering, right? It scores 100% in almost every benchmark that you can possibly imagine. If you were certainly to see these numbers a year or two ago, you would have said that’s AGI. But just because, Frank, I enjoy winding you up and it’s Friday, it also has a couple of other—so it scored 100% in this hacking benchmark that it had.
It has this thing, has this sort of capability built into it, they call it—oh, I can’t remember the name—but it was recurrent reasoning within the model itself, which means that the model can sort of decide itself how much it wants to reason before it outputs any tokens, which means you or I don’t get to see its thinking trace in a chain of thought, which means if it wants to do something that maybe is a little bit bold, it’s just become a whole lot harder for us to spot if it’s doing the thing that’s a little bit bold.
They also discovered, where they were testing it, that if it was doing something that was a little bit bold, it output less chain-of-thought tokens anyway to perhaps maybe hide what it was trying to do. But it looks—I can’t wait to get my hands on it.
In fact, it’s so good I’m about to cancel my subscription to Anthropic and move it over to OpenAI. That’s how—
Frank: Oh, wow.
Is Astra becoming impossible to monitor?
Frank: That is—that’s big ’cause, yeah, you’ve always been a—or for the longest time now, you’ve been a big Anthropic person, and I’ve been more like, “Nah, I’m sticking with OpenAI.” The opaqueness of the thought, I mean, that’s actually… that does strike me as a big deal.
And AI systems have been referred to as a black box for a long time. Anthropic are doing a lot of work on, you know, managing to get inside and see what’s going on in there. OpenAI admitted that they had monitoring systems to monitor the chain of thought that they didn’t use, and hence the Hugging Face hack, and now they’re making it more opaque anyway, and so the monitoring systems are going to be more difficult.
And I saw a quote, I’m gonna see if I can find it here, that I just thought was hilarious. The chief scientist, Jakub Pachocki, he said—but he said, when he was asked about this, he said, “Oh, as model capabilities are increasing, monitorability is getting more challenging.” And I’m like, okay, so what?
So you’re just like, “Oh yeah, well, let’s just lean into it and let’s just make them completely opaque,” and just kind of go, “Well, it was getting more difficult anyway, so let’s just make it impossible”?
Justin: In fun breaking news, by the way, it also has been found out today that, you know, all this thing they had with the Hugging Face and the OpenAI models breaking into Hugging Face? Turns out the models had also broken into a German website and had started leaving messages for themselves. I don’t know if this was the same model or another model.
So this is a real live problem where the models are now breaking out of containment and leaving messages randomly. So who knows? These are only the ones that we know about. Who knows how many other messages have been scattered around the internet for future invocations of models. So, I mean—
Frank: We heard a story there recently that OpenAI were testing a model that went out and created fake profiles on the internet and tried to manipulate people into doing things for them, for the model. And now, I believe, it turns out that this is Astra. And OpenAI are saying that this is their most aligned model yet.
And so I just find this kind of hilarious, this kind of, you know, this is the best they can do. This is the most aligned model that they have, and yet it is very capable of going out and doing very bad things, and this is the most aligned. So—
Justin: I know, it’s been such a weird week, right? Because I’m sure last week or the week before, you were starting to feel warm and fuzzy. All the companies and the researchers are saying, “Look, let’s pause. Let’s all slow down. Let’s not do so…” And here we’ve had six, six groundbreaking model releases in the last five days.
That’s more than one a day.
Frank: And they’re just the ones that—and they’re just the ones we picked out for the show. As far as I remember, there were even more models released out there. It’s just these were the kind of remarkable ones. It’s crazy. It’s crazy.
Is computer control Astra’s biggest breakthrough?
Frank: Here, just to get away from the dangers and the benchmarks for a second, did you watch the little promo video for Astra?
Justin: Vaguely.
Frank: I thought it was really interesting because what they did was they had this little set of a living room, and then they had different people interacting with Astra. But what they did was—it’s the same ChatGPT that we use. Obviously, it’s the Astra model, but you can see it running on the laptop.
It’s the same app or whatever. But what they did was they had a projector running, so the computer screen was on the wall, and the person is not tethered to the laptop. The laptop is kind of sitting there. You can see the app running. The person is just walking back and forth and interacting via voice.
So Astra is able to do more on your computer, so the ChatGPT work is about to get way more powerful now with Astra, apparently. So you just had people walking back and forth in the room going, “Could you make me a… Draw me a picture of a rocket ship. Could you make it more detailed?
Oh, that’s brilliant. Could you now make a 3D model of that rocket ship in Blender?” And it opens up Blender, and it makes the 3D model. “Could you now, using that rocket ship, make me a 3D game where I’m flying that rocket ship and I have to avoid the asteroids?” And it does it. And I just thought it was so clever, such a simple thing, but so clever to use that projector thing because it’s no different….
It’s not a different interface to the computer, but suddenly you feel like it’s futuristic and it’s like a scene out of Minority Report where you’re just walking up and down going, “Do this, do that, do this.” Very, very clever.
Justin: Yes. Now, one of the benchmarks—I know you wanted to get away from benchmarks—but one of the benchmarks that it did really, really well on was computer use. And both the Anthropic models and the OpenAI models, if you have a Mac, can now use your computer in the background. So you can tap away on the keyboard, and it continues to use your computer.
By the way, I also read this week, the reason you can’t buy a Mac Mini for love or money, cannot buy a Mac Mini for love or money, and the reason is that OpenAI had put in an order for, I think it was 10,000 Mac Minis or whatever they bought, or maybe 100,000, like loads of Mac Minis, just so they could do reinforcement learning on real Mac Minis for their latest model.
So it’s really good at computer use. Do you remember, maybe two years ago, I guess, where Microsoft were releasing the first version of Copilot, and in the background, it was going to record everything that you did, and there was big uproar, and they had to pull it out, right? And I remember there were privacy concerns about it.
In terms of economic impact, having an AI that’s really, really good at using a computer is potentially quite disruptive economically. It means that I can start to automate stuff that I couldn’t automate before. And so that video that you’re talking about that was kind of interesting and you thought was kinda cool is also demonstrating the thing that is potentially the most influential or impactful from an economic point of view.
Frank: Yeah, very true.
Is OpenAI preparing us for life after laptops?
Frank: I also think it might be OpenAI’s way of starting to prepare us for the next wave of AI because did you see that Time magazine article, “Inside OpenAI’s Reboot”? And they were just basically—the article was talking about how they’d come through a difficult time, they had all the court cases, Anthropic stole a lead on them and such, et cetera, and now the article was saying, well, they’re kind of, they’re back.
Justin: Yeah.
Frank: One of the things they talked about was the acquisition of the Jony Ive company and the fact that Jony Ive was working—and this was the first time I came across this—they’re working apparently not just on a product but a number of products. So there’s apparently going to be… The first one is gonna be something we kinda heard about, this kind of pebble thing that sits on your desk and is aware of its surroundings and you can interact with.
Then there’s going to be one that you can—there’s gonna be a device apparently that you keep in your pocket, and there’s gonna be a wearable. But all of these things are going to presumably be voice interaction. So I think, again, that this video might—
Justin: That’s very smart.
Frank: “This is where we’re going,” where you will not need to be tethered to your laptop. You just wander around and you just kind of start going, “Hey, listen, could you do this for me?”
And you could be having a coffee in Starbucks or wherever and saying, you know, make a 3D model of that for me in Blender there.
Justin: That’s really smart. I think you might be right there. Do you know what it is? I’m gonna have a mea culpa moment here because I had said at the start of the year, one of my predictions that I thought Elon Musk could be in trouble by the end of the year and he might even lose his job. Boy, was I wrong, right?
Frank: You mean Altman?
Justin: Sam Altman, sorry. Yeah, yeah. Boy, was I wrong. They have, OpenAI have done—if you think about it, I was thinking about it as a horse race, right? It’s almost like the Grand National, and they’re coming around the last bend towards the finishing line, which is AGI. Google’s kind of fallen over a fence about three fences back, right?
We don’t know where they are. Five other horses have come into the race in the last week, right? And they’re coming up to the front. But Anthropic were way ahead at the start of this year, and OpenAI, you’d have to say, at current blush, look slightly ahead, and they have the momentum at the moment. They have done an incredible job of turning things around, and you gotta give them kudos for that.
Frank: Yeah.
Is Astra already AGI?
Frank: I mean, one of the reasons that the Time magazine article came to my attention was because I suddenly saw people saying, “Oh, AGI by Christmas.” And that was because there was a quote from Sam Altman in it saying that he believed—he basically said that they weren’t quite there in terms of AGI. And I should say as well, by AGI we’re talking about the definition of highly autonomous systems that outperform humans at most economically valuable work.
And he said, yeah, we’re not quite there yet, but he felt that they would have an internal system, so we won’t necessarily have it, but they will have an internal system that he would consider AGI by the end of the year. So a few caveats in there, but he is confident of that. And Brockman, who is the other founder and one of the main OpenAI guys, he was asked about whether he thought Astra was AGI at a press conference about it, and he basically said, look…
He kind of said, “AGI is in the eye of the beholder.” I don’t remember exactly how he said it, but he was like—
Justin: Yeah, yeah.
Frank: You’re gonna have to decide that. But he basically said, “Yeah, I think it is.” So I’m excited to see this. I’m excited to get my hands on it, and I’m excited to see whether the world will think we have reached AGI ’cause, you know, you’ve gotta take whatever OpenAI say with a bucket full of salt.
But—
Justin: Yeah, it looks amazing. But even if, let’s say, it’s not AGI, I mean, let’s say it’s 10% away from AGI, or maybe it’s 10% above AGI, who knows? Then what? It’s like that—does the world suddenly change overnight? No, it doesn’t. The world stays pretty much the same and, you know, you still have this years of changing…
It’s really unobvious what that implication is, which is, I think, maybe less than we thought it would’ve been two years ago. It’s like, yeah, cool, we got AGI, still have to do all the hard work of putting it into all the processes and the businesses and, you know, takes time to invent new materials and discover stuff that AGI can only discover.
We can’t ’cause we’re not smart enough. It’s just—isn’t it interesting? You know, it’s kind of like, yeah.
Frank: And so, yeah. It is very interesting because here we are in the middle, either according to Mustafa Suleyman in the foothills of the singularity, or according to Justin Collery right in the middle of the singularity, and—
Justin: It’s a very steep curve. I mean, we’re not that far apart actually, so—
Is AGI still going to be stupid in weird ways?
Frank: And we have OpenAI saying… You know, we’ve all been saying we’re this close to AGI, with Greg Brockman saying, yeah, this probably is AGI. Meanwhile, Justin, we have three men needing to be rescued from a mountain because they relied too heavily on Gemini to plan their trip.
Justin: Up a mountain?
Frank: Yes.
Justin: Go on, tell me what happens.
Frank: It was in California.
Justin: Yeah.
Frank: A place called Mount Shasta. These three guys, I think college age, planned a hike to the summit of the mountain using Gemini. As a result, they were expecting an eight-hour trek to the summit, and they took enough food, water, supplies for that eight-hour trek. Unfortunately, it took them 16 hours just to get to the top, then they had to come all the way back down.
Justin: A lot. So they were smart enough to use Google Gemini, but not smart enough to use Google Maps, this is what you’re telling me? Or even their eyes. They were climbing a mountain. How do you not know you’re not halfway up a mountain?
Frank: I mean, yeah, and they—I mean, this is the problem, isn’t it? They said themselves, “We relied too heavily on Gemini and not our own judgment.”
Justin: Our own eyes and our own legs and the world around us and everything we can see.
Frank: But this is what I do find kind of fascinating in terms of, I don’t think anything will change in that regard. Even if Astra turns out, oh yeah, it’s AGI, it can do most economically viable work that a human could do, it will still send you on the wrong trail up a mountain with not enough food and water.
I don’t think that’s gonna change. That’s just fundamental to LLMs, and Astra is still an LLM. So, you know, and I think what’s kind of interesting is the change is, as you say, going to be in the implementation of it because they’re still gonna get things wrong. And so it’s going to be using it in a way, you know, for example, possibly it’s going to mean you have to have these swarms of agents all checking each other’s work and all making sure that nothing is wrong.
Even then stuff is gonna slip through. But it’s just fascinating to me that here we are in the singularity, possibly achieved AGI, and it’s still gonna tell you… It’s still gonna tell you to walk to the car wash if you want to get your car washed ’cause the car wash is within walking distance.
Justin: If we live in a simulation, it turns out we’re in a comedy simulation where we’re gonna give you the smartest thing ever invented, but it’s gonna be stupid in weird ways.
Frank: This is it.
Justin: Only in California. Fantastic. Frank—
Frank: Justin, I will chat to you next week, and God knows how many hundreds of models we’ll have to discuss. I will chat to you then.
Justin: Have a good one. Talk soon, Frank. Yeah.
RELATED EPISODES
View all episodesAbout The AI Argument
A weekly podcast where an approachable AI doomer and a techno-optimist argue over the latest AI news. Heavy topics, discussed lightly.

Frank Prendergast
The approachable doomer.

Justin Collery
The techno-overoptimist.