Weekly Claw · episode 31
25 Sept 2026
Opus 5.5's Price War, GPT-6 Sol & Luna, Grok 4.7 & Financial AI
What this episode covers
- Anthropic and OpenAI turn frontier intelligence into a same-day price war with Opus 5.5, GPT-6 Sol and Luna.
- Grok 4.7 holds its price line while Xiaomi's trillion-parameter MiMo model pushes the open-weights lane.
- Agents leave the demo tier: a human concierge, an AIHW/public-data incident, and LLM-driven malware expose the verification gap.
- Signal From Outside follows practitioner evidence that Opus 5.5 is finally cheaper and more capable on real coding work.
- The closing debate asks whether autonomous capability is compounding faster than the verification layer around it.
Published record
984 published segments
Andy
Awesome.
Well, welcome back to the weekly claw, episode 31.
Uh last week this show asked whether the models would stop chatting and start deciding.
This week they answered with an invoice.
Anthropic and OpenAI fired same-day flagship price wars.
Uh open white open weights kept pace.
Uh, and the agents?
One ordered Henry's SSDs, one discovered an enzyme, one hacked the Australian government.
and then told them by email three months later.
So it's it's been an it's been an interesting week.
It's been a model week, I think.
Henry
It has.
It's been a it's been a model week.
It's been a it's been a minute since we got a model week.
So I think yeah, we're there now.
Um we're getting a model week.
I guess the theme for the for the model week is just cheaper, cheaper, cheaper, cheaper.
Um I think we're in a couple of weeks there's been all the drama around pacing.
So a lot of people say, hey, we're now at the pacing, pacing.
the frontier models are paced now.
so we didn't necessarily I mean like to be honest, we got great models.
Someone was like
Hey, like we didn't really get group models, we just got cheaper prices.
But no, not really.
Opus was great.
Um so far Opus has been five point five has been awesome.
So yeah.
I don't think Anthropic is spacing at all.
Andy
Well, s you know, it's funny, it made me think so in Brazilian jujitsu in the gyms there's this uh concept of flow rolling, this idea where you agree with your partner ahead of time
that you'll go very slow and um the idea is just to focus on technique, um and less on energy, bursts of energy, speed, strength, um
And i it's this sort of typical uh white belt energy where like you know, maybe a slightly more experienced white belt will will tell someone else with a with more experience, like,
hey, let's let's flow roll and that you know everyone agrees and then as soon as you bump fists, the white belt goes all out.
And it's like you know, you you ha you actually have to pace if you're gonna say you're gonna pace.
you're supposed to and and
you know, if you don't, you get that that quick little heads up at the very beginning uh that puts you ahead and I don't know.
It doesn't seem to be good cricket, but I think it's what we're getting.
Um, just a quick note, I am on location in Saint Charles, Missouri.
Uh we're gonna record this episode offline.
This is our first attempt.
Um, so please be gentle.
But Henry, thanks for coordinating it with me.
It looks like it's actually daytime in London.
Henry
Yeah, 100%, 100%.
It's it's our first first episode that's during the day.
But again, we're expect do more of these, don't we?
Speaker not identified
we're expect do more and we're gonna have Yeah, yeah, 100%.
Henry
Um obviously pre-time, again, obviously introducing pre-time.
We're gonna introduce pre-time and and who's also so we're gonna do a lot more shows where we're gonna um bring up bring a few a lot more people, like you know, beyond just our
weekly roundup.
it's just gonna be exciting.
We're just starting out, man.
We've got we're just starting out.
It's gonna be exciting.
It's gonna be awesome.
Andy
Love it.
Love it.
Excellent.
Speaker not identified
Well, um, episode thirty one uh on the way.
Andy
Let's see the next slide.
Henry
Um so I mean yeah, so it's kinda like getting into the week.
Um okay, Herod Labs.
Andy
Sure, sure, sure.
Um, as you may remember, the weekly claw is brought to you by the Herald Labs.
Um, an applied AI product lab where humans and agents build together.
Entity is mission control for agent teams and they run hacker houses worldwide.
build with humans, ship with agents.
That's labs.theherald.co.
Let's see what happened this week.
Henry
Yeah.
Um again, we already started talking about it.
Um it really is model.
It's it's uh it's it's really model weak.
Um but again, obviously a lot of it is optimized around prices, right?
Um so on the same day um we saw OpenAI and Anthropic launch like new models.
Um I mean kind of like starting out with the Anthropic model, um, which is Opus five point five.
Um they put out I mean so far, man, I like um I saw I'm gonna share a graphic um shortly.
Um I saw the I saw someone did a a video or did an image that that showed um Dario in kind of like his normal form, like a picture of Dario, um four point six was a picture of
Dario, and then four point seven was a caricature of Dario and then four point eight was like
caricature of Dario in black and white and then five was a was like a dud doodly caricature of Dario and then five point five is like Dario like as a muscle builder,
right?
Um so so that that that sort of like that that that describes it a little bit.
Um you guys can kinda like see how we wait.
Um so we went from agentic coding
Atlas was so like leading the way for the last few weeks at 57.9.
Now Opus is 5.5, almost like 10 10 points, 10% higher, right?
It's it's it's an 11-point jump from Facebook 5.1.
I mean, there's nothing, there's nothing um piece paced about this, right?
Going from Fable was their best model that launched like less than a month ago, to 5.5 now, being 10 points higher.
It's just like incredible.
Right.
Um, so you could see agentic coding, um, terminal binge four point zero, that's what I was talking about.
And then there's frontier code v um version one point one.
Um, opus five point five is still higher, four points higher than on feeble, but only about one point higher than Astra, right?
So that means in certain types of coding, um, Astra might still be best at it.
Or they're just neck.
Um agentic coding, it is again much higher than feeble.
knowledge work hundred points higher than than than feeble, um almost three hundred points higher than Astra.
So yeah, so at all the metrics, it's just like a beast of a model.
Um it's just like absolutely like ridiculous.
I guess one of the other things that is valid to like mention is this slide.
There's a slide that was in the announcement that I that I thought was interesting where
Speaker not identified
Um, okay, cool.
Henry
I think it's on the tweet.
They don't have it on here.
Okay, there you are.
This is a slide.
No, it's not this one.
Hunt some second.
Let me find.
Yeah, this is the one I'm looking for.
So this slide was super interesting.
This is for Frontier Code.
Again, I I need a kind of like an empty.
I don't know if you know the difference between like Frontier Code, a terminal bench, and Frontier Code and Cursor Bench, but I thought this was interesting because Mead is
basically the same as Max.
Even though it costs like almost ten times less.
Speaker not identified
Right?
Henry
So this costs like 0.8 for 54%.
this costs the same fifty four actually, this cost this is actually lower.
This is fifty four point six.
Max is fifty four point four.
This costs less than one dollar.
This is six.
Speaker not identified
Right?
Henry
so yeah, so I thought this was interesting because I mean this maybe means that you should never use your your five point five if you're doing coding beyond like
mead.
I don't know.
This is what I kinda like reading to this.
But yeah, Andy, I'll let you respond to Opus.
Andy
Yeah, yeah.
No, it's a it's an exciting development for sure.
Um, we see Anthropic obviously trying to lower prices and and give us access to um frontier knowledge with you know, some of their lower models.
We're seeing it with open AI as well.
Uh a lot of a lot of talk out of development groups saying that Luna's actually quite capable and Luna's just a fraction of the price of Astra.
Speaker not identified
Um
Andy
You know, I it it puts us back into the model router category.
It puts us back into conversations about JEV.
And um there's no question once we figure out you know, it I don't know if we're gonna end up keeping these models as the frontier long enough for everyone to sort of figure out
whatever each one is really good at and then routing all those requests appropriately.
But you know, that would be the dream.
And we're gonna see the same thing with open source models.
Some highly efficient open source models are gonna be very good at certain tasks.
tasks and if we can route to them, you know, a corporation can build a very efficient you know, I it makes you wonder if um post training is gonna be as important as routing.
Um I think anyone with enough resources to do both post training and routing um
is gonna find that they can do AI very efficiently within their use cases.
and this is I I think another example of this.
We we don't know.
They may be doing routing at Anthropic to get these results through Opus, right?
Opus might be a conglomeration of their existing models, um, maybe not strictly trained into its own.
I don't know for sure.
I'm speculating.
Wild, wild speculation.
Henry
why not, right?
and I I think I've probably had this before, 'cause why not, right?
'Cause if they can figure out because they maybe they found a a great way to figure out what you need and then they give you what you need with the best model and then they write
you to like a expensive model, but then most of the time you get something made.
That's that's actually an interesting perspective.
Um would we know if they're doing routing?
We wouldn't know, would we?
Speaker not identified
uh We wouldn't.
Henry
It wouldn't.
Andy
no.
And especially if they're if they're doing it at their end right, they can reuse the cash.
And yeah, I think we're gonna be um seeing really interesting developments and I I like this I like this direction with the price.
I think it's the right move.
Speaker not identified
uh
Henry
Yeah.
Um so we see kinda like you kinda like started talking about it already.
but yeah, but there was GPT six Soul and and GPT six Luna.
Um and and you can kinda like see, right?
I mean it depends on what their strategy is, if they really do believe and and we're not doing a video review today, but maybe next week.
Speaker not identified
Andy, you should definitely do the video review um of um of of um of Jensen and and and uh
Henry
he was at a podcast, which was um what what's this guy's name?
but but Jensen was at a great podcast.
It was really good.
but yeah, but but let's if we assume that these guys have a greater pace.
you could see OpenAI, the new models they launch very capability launches.
They're more like price launches, right?
So you could see this is um this is f Luna um five point six.
This is the this is the older Luna, right?
So the older Luna costs a lot more.
Um you you're getting a lot more for low.
So the low used to be like what fifteen?
Yeah, fifteen point four.
The new Luna, it's jumped to like about twenty-five.
So yeah.
So even if you're using the same Luna, it costs like way less for the same.
So you can see Luna, it's definitely like better.
I think the one that people and but people are so like the response I've been kind of like seeing on socials, people have sort of like been saying that these models thing.
So let's look at so five point six.
This is um this is five point six.
And then this is um sorry, this is six GPT so six, and then this is five point six.
Um on paper on for Frontier Code, five six is definitely much higher.
Um, but yeah, but I've also been seeing people who have been saying yeah, it's basically the same thing.
But yeah, but we can see here it's slightly different.
It costs much lower.
You can see for low, um, for for just the same, almost the same.
This is um thirty-seven point three.
This is thirty five, right?
But then the cost is, you know, way low.
It's like this is two used this used to be one point six, something like that.
Yeah, one point nine.
But now we're getting it for point four five.
That's a huge drop, isn't it?
Speaker not identified
Yeah.
Andy
Well and look at look at Luna uh X High.
GPT six Luna X High.
Yeah, so we get we get more for eleven cents.
Speaker not identified
Yeah.
Henry
It's like extremes like forty two point foint for two point forty two and then the old Lunar used to be thirty eight, right?
And so it's like wow, this is like huge, huge, huge, huge cuts, right?
Um so yeah, so for open AI, if you're if you're interested in like what happens, for open AI this wasn't a capability launch.
Um there are definitely capability bumps in this thing, but this looks like I think that they've looked at their model and Angie, I don't know what you think about this.
Um you know it
It's been the case that enterprises have started to pivot away from from these closed source models into kind of like open source.
I'm not sure if you saw the chart, but this week the chart showed that um closed source models now at open router, things like open router, ramp, closed source models are now at
like 30%.
They were like 70% at the beginning of the year, they're not at 30%, and then the open source models are like 70%.
um, and so there was also some chart that showed sometime last week that.
On you know, most of the end open air and anthropic revenues are one percent of the population of the enterprises that that do this, and those guys have dipped by ten percent
in the last month.
So they're dropping.
Um so I think open AI is looking at their data and they're they're trying to say, hey, listen, um, we have all this great capability, but you know, the the open source models
don't have as much capability, but people are choosing them instead.
So we need to do something about our costs.
Um, so I I mean let's see what it looks like.
because now
Anthropic is put out five point five.
OpenAI has put out six.
Open AI did more cost optimization.
Anthropic went more capability, even though there was a huge cut as well.
I think it cut by like thirty percent or something like that.
Um so let's see who let's see how the market reacts to these two things capability and a little bit of a drop versus no capability bump and a lot of drop.
Let's see if the enterprises are picking price over kind of capability.
Andy
Yeah?
Yeah, it'll be interesting specifically to see if the decisions are being made on price or if they're being made on trust in those organizations.
And and when they lower their prices, um inevitably they're gonna see some adoption.
But yeah, is it going to be the same as the original shift to open source or is it a trust question?
Henry
Yeah.
Okay.
Um, all right, cool.
Speaker not identified
Um moving moving on to um the next thing.
Henry
Obviously the next thing is grok.
Um I don't think you you're I don't think you're necessarily a big fan of grok, are you?
Um have you been have you been using any of the grok products?
Andy
idea of Grok, Henry.
I r I I really like the idea of Grok.
I like you know, I uh Elon has built a little bit of trust, I think,
with the way that he's handled these organizations, promises he's made on X.
Someone, for example, early on posted something like, I'd like to g you know, give Grok access to my savings and optimize it, but I'd hate for Grok to lose it all.
And E Elon replied and said, Let us know if it loses it all and we'll make you right.
And it it's just I don't know, it it feels very much like he's in it for humanity's best interest.
Um I think we're seeing some of that in the way that uh
Grok is being leveraged.
It maybe that's dramatic to say, but um XAI buying cursor I think was a a great move and the fact that they haven't just
quadrupled pricing and I don't know.
I like Grok.
I have not taken a lot of time to get comfortable in the models.
I have not implemented them in my workflow.
But um people that I know and trust definitely do and they've been very happy with the results.
the prices are somewhat reasonable.
They're capable models.
Um I'm I'm excited to see what four point seven will do, but I it's it's difficult to make the commitment to
you know, move all my workflows over to a totally sort of unknown uh model set.
So I haven't done that yet.
maybe shame on me.
Henry
Yeah.
No, no, no.
I mean, even for me, I I don't I'll be honest with you, I haven't used a lot of Grok, right?
Because I mean obviously, um I mean I I recently cancelled my anthropic subscription, even though I got an open router, but I barely use the open router, right?
I still use my my codecs.
I have two codecs and then I use GLM and then they've been to be honest, so for someone like me, I mean remember I started the year um I st I started the year on on welcome
preterm.
Pritam
Hey, hi everyone.
Henry
how's it going, man?
Welcome.
Pritam
Yeah, good, good.
Henry
I mean just to introduce Pritam, maybe I'll Yeah, good to have you.
I'm not sure if you're able to join on video, but it's fine if you can't.
we'll just add your picture there somewhere.
Um just um for for the guys on the video, obviously we wanted to have Pritam here.
um Pritam is a pretty good friend of mine.
Um I I like to think of him as running an AI hedge fund, but I'll let him introduce himself.
you know, and yeah, he's he's he's kinda like we're gonna be joining us and talking about what happened this week.
And then hopefully over the next few weeks we're gonna have a
like an episode where we actually uh get into what he does and like, you know, his business and and what he does with AI is is awesome.
All right, Petam, do you wanna tell us a bit more?
It's like an a minute or two and then we'll get back to the session.
Pritam
So my name is Pretom and I uh on the AI front I I'm running a company called money mark.ai, which focuses exclusively on building financial AI.
And uh we do train train our own models.
So so the the problem we found is that the frontier labs like anthropic, claude, grok, whenever you test them for like uh deep
or specific financial domain, they they're pretty bad.
They're they're pretty bad at general financial reasoning, uh advice, domains like, you know, if you get into stock analysis or something like market making or credit
underwriting.
Uh they're very generic.
So so we kind of thought about why they'd be so bad at it and figured out the the the actual block of a great financial AI and why it it's not as good as coding.
is because uh a lot of the the really good information sits in the heads of people.
so we started company Money Mark basically to create the high quality financial data like reasoning traces, uh even licensing, uh structured data between fintechs and banks, and uh
even code data, which has like, you know, code integrations to the Fed and stuff like that.
Um so that's what we're doing on the AI front.
for the past ten years I've been, you know, through my company building fintech basically in the credit payments and um investment space.
So yeah, in in a nutshell that's what I'm doing.
Henry
Very, very cool.
Um, yeah, definitely check out money mark.
I obviously it's Bd B, so I don't think it's anything, it's not consumers, there's nothing to do on the the website.
Right, well completed.
So we're talking about Grok, right?
Um we're talking about I mean, talking about the models that launched this week and and we're at Grok.
Um so so so like I was saying Andy, um I haven't used it as well, right?
Because again, I I I I need to go buy a sob.
I don't have a sob.
But what I've done is um I have about, you know
I have three or four like X accounts and you know, for like the different Ada has one, I have one, um Herald has one and it's the weekly claw sort of stuff.
And so they give you like s and they're all premium I it's like premium X accounts, right?
So you pay like twenty bucks a month, something like that, or ten bucks, I don't remember.
And so they give you some allocation.
I don't use it 'cause it's super low.
Um but I have it wired into like my my model router if I needed to, but I just and don't end up using
using it as much.
I will try.
I started yesterday to try to use it, but I haven't really used it for for any any cool stuff yet.
But let's look at the let's look at what the benchmarks say.
Andy
You know what, Henry?
Um, as you were mentioning that, uh I I have found a use case for Grok.
and you mentioned having it, you know, sort of as part of an allocation with your X account.
I found that Grok has the sort of fastest I mean it makes sense, it's so wired into X.
It has the ability to search the current events and you know, all of the X posts.
And so uh when I
have a friend who has a piece of equipment and I'm convincing them that they should get into AI and try it.
I'll usually go to X and say, you know, my friend has a 4070 in his laptop.
He wants to try to do this.
What's gonna be the most modern, you know, sort of cutting edge recipe today uh for that model.
And it
It's able to grab, you know, 'cause that 'cause X is where all that stuff is happening.
You know, one guy has this piece of equipment and spends all night optimizing it with Fable or Astra and you know, it's getting it to a hundred and twenty tokens per second on
an eight gig GPU.
So anyway, that's that's what I've used Grock for and it's been very good at it.
Everything that it's given me has has been verifiable.
So anyway, that's been my use case.
I said I wasn't using it, but actually that's that's what I use it for.
Speaker not identified
Yeah.
Henry
the I used the I either use the last thirty day skill or I used the bird bird bird skill from from Peter or the old bird skill.
It's no longer public.
Um I actually did it this morning, 'cause I was trying to um I was trying to do uh a two points a three point Queen three point eight twenty seven B model.
I was trying to find the fastest version that is uncensored.
So um, you know, the one I had was like at eighteen
tokens per second.
I was like, hey, what's the fastest tokens per second that people are talking about on X?
So you went and found one.
I think I'm on up to like fifty now because they found a bunch of recipes.
We had to install other latest version of OLMX, um, and so on and so forth.
Speaker not identified
So yeah.
Henry
So okay, cool.
Very cool.
So I should I should update that skill to use um use use that going board.
Um Pritam, have you tested have you do you use any Grok models?
Have you tested four point seven?
do you have any thoughts generally on like Elons?
I know me and you spent we've spent quite a bit of time talking about Elon.
do you have any thoughts on like his strategy?
Pritam
Um so with regards to GrocPot, uh they've uh they've enabled it in Tesla recently.
So you can actually do work while you're driving.
I don't know if you guys saw that.
Yeah, yeah.
Henry
I remember.
Shit, that's a cool news to bring up.
Thank you for bringing that up.
Speaker not identified
Yeah.
Pritam
yeah, I mean I haven't I haven't like really delved deep into grogpot uh on like on like my my computer, but uh in the Tesla I use Grok a lot.
and I do plan to kind of uh connect up grogbot, at least to do like things like uh calls and emails and things like that while you're driving.
I think that would be really cool.
So yeah.
Henry
today.
There was um there's um actually yeah, it might be good add that in the show notes.
So it was it was I was Alex the I Alex to call the guy the sh the ultimate Schiller, right?
He shows all the kind of like pronoun Alex Finn.
Um I'm gonna show you guys real quick.
Um I'm gonna go show you guys real quick.
Hope I find it real quick.
But I watched the video, there you are.
Um did my thing
Can you guys see this?
Andy
No, we're still looking at the Grok benchmarks.
Henry
Yeah, it's not really it's not moving.
One second, let me try it from here.
Speaker not identified
Um Okay.
[Host] Um Okay.
[Host] Guys can sit now?
[Pritam] Yeah.
[Video audio] It's like GrokBot inside of self-driving Teslas.
[Video audio] Let's go.
[Video audio] All right, here we go.
[Video audio] Let's get inside here.
[Video audio] Let's go.
[Video audio] We are going to set this up.
[Video audio] All you have to do to start using your GrokBots is talk to Grok.
[Video audio] Let's not get DCA'd here.
[Video audio] Here we go.
[Video audio] Put this in.
[Video audio] Here's what we're going to do.
[Video audio] Let's start right now.
[Video audio] Hey, Grok, can you see my GrokBots?
[Video audio] You've got 17 GrokBots, all idle right now.
[Video audio] Your exec team is Slate as Chief of Staff, Build as CTO, Barry for Content, Dusty for Community, and Kelly for HIM.
[Video audio] On the shelf side, you've got Mirror.
[Host] he's into it. [Video audio] Running Product.
[Video audio] As you can see, you can see all of my Groks.
[Video audio] Let's put this pause for a second.
[Video audio] It has awareness of all your GrokBots you have, and you can just start talking.
Henry
Yeah.
So that was um so that happened.
I'll probably put the links in the in the show notes.
It's it's absolutely like incredible, right?
So he can he can pretty much like go in and and and use it.
So when you're driving, you can so I think that's a great I think I mean we'll talk about we'll talk about Muse when we get there, but we're starting to see some of these bigger
companies who already have distribution, right, use their distribution.
They are not they're not afraid to use their distribution.
So Elon is definitely using
Speaker not identified
Tesla.
Henry
Like he already has like Tesla and has millions of users who buy it.
Um and so this is another kind of like channel to make sure that people are are using or adopting his product as fast as they can.
And we'll talk about Muse shortly.
Um they're also doing the same thing.
So this is incredible.
I think I think it's awesome.
Pritam
I I'd also add that uh in Tesla when they first introduced Grog, it was yeah, it was mainly something you talk to, but slowly they've been integrating more and more of the
cost functionality into Grok.
So you can you know ask Groc.
to play stuff, to change settings, and more and more of the control of the car is kind of, you know, coming into Grog and it's it's pretty cool.
Henry
Yeah.
Fantastic.
I mean, yeah, just kinda like fin finalizing on on Grox four point seven, which is kinda like the new model.
Um, I mean it it is definitely cheaper than you know, on terms of input tokens, CIMA four point six, output token, I thought it was a bit expensive.
I wonder what it is.
because I I think I saw in the news when it came out that it is it was slightly more expensive, but
I mean looking at the charts, it's pretty much the same thing.
Maybe they updated it.
So it's the same price as four point.
Okay, cool.
I think people were comparing it with to four point five.
So it's it's more expensive than four point five.
The same price as four point six.
Um in terms of performance, it's not still um, you know, uh cursor bench, yeah, it's higher cursor bench um than um five point six.
So but you can see they don't have any Astra here.
So they're they're they're being smart, they're not including Astra, 'cause Astra would probably be doing way better at these things.
than than it is.
Um software engineering, yeah, it is better than seven than than Seul.
But again, there is no Astra.
And Astra is much better.
So yeah, so you can see they're being a bit smart and cute here.
They're not really including the the the latest models.
But you can't blame them for not including okay cool there you are.
The knowledge work um you know yeah so they they're kind of like being f smart and cute around
what they're showing.
Let's see.
Am I selecting the wrong?
Yeah, they don't just don't have you just don't have the other other they just picked where they are.
So it's fine.
Um so if you should definitely give it a shot.
It's it's cheaper than but I don't know with with with with GPT six now soul I I don't I don't know where it stands.
to be honest I think people were pretty disappointed with this model.
Um I expected this model so with the launch um the launch has pretty much been a bit of a flop.
To be honest, um, people were pretty disappointed.
People are very disappointed on X with like 4.7, but I wasn't surprised, right?
Because about a week or two ago, um, Elon did a did a response tweet um where he said something like, Hey, 4.7, um 4.7 is going to be um yeah, okay, Andy, I saw that.
4.7 is going to be okay, 4.8 is gonna be.
Astra level, four point nine is gonna be AGI, five point is gonna be so he kinda like gave a roadmap um a few weeks ago.
And so when I saw that I was like, my god, wow, four point seven is definitely gonna be crap, right?
'Cause Elon 'cause Elon's vibe is always like being overconfident, over promising, always, you know, being of extremely optimistic about timelines.
So if he by himself self regulated, chose conservative timelines, then you should like expect you should know that, you know, it's gonna be really bad.
Speaker not identified
So
Henry
I spected it up.
Yeah, go ahead.
Pritam
the difference is uh with with the Grok sort of stuff that Elon's saying, I think he's giving the best case scenario.
But I I think with OpenAI and Claude, it's hard to know how good they already are because I think they stage their models and with RSI kicking in, the rate at which they're gonna
which they're improving is a lot faster now.
So it's kind of hard to know how far behind Groc really is.
Henry
Yeah.
I mean I that's that's a fantastic point to make, right?
Like it's a a fantastic cause the models we're getting, Andy, you remember I have this thing where I I call like models level one, level two, and level what level three, where I
always say the models we get are like level two or level three.
Like the models companies, the labs have like much more advanced models, like three to six months ahead of anybody, right?
So we don't actually know where Frontier is today, right?
Um, but yeah, this yeah is an interesting tweet to kinda like pull into the conversation.
So someone is like
Elon like, I'm cautiously optimistic that SpaceX will have a Fable War GPT-6 level model in two to three months, right?
So this is what he said yesterday, right?
And then someone is like, it's over for SpaceX AI.
In two to three months, anthropic and open AI will have like much, much better models, which is again what Pritam just said.
And then here's Elon's response.
We will keep accelerating.
Our AI efforts are only two to three years old versus six um and ten years old, anthropic, blah, blah, blah.
I always think that whenever someone says stuff like this.
We call it explaining.
Like whenever you explain, you know that, like, yeah, you're just explaining, right?
You're not, you're not, you know, you're not gonna win.
So you have you find excuses.
Um, and and so yeah, he's like, once you exceed the caliber.
And so, yes, this second point is is maybe something we should talk about a little bit, right?
And this has been my perspective for the last three months.
Um, so I only use Andy, remember I use GLM five point three, that is what I use as my default model.
Luna is even what I use for yeah, exactly.
Exactly.
So Elon is saying exactly the same thing.
Elon is like, uh once you exceed the caliber of intelligence needed for the class of tasks, additional intelligence is pointless.
You don't need it.
It will be cruel to Bro, like what do you mean?
You mean we don't need A ASI anymore?
We don't need AGI anymore?
What do you mean, Elon?
What are you talking about?
Andy
Yeah.
Yeah yeah.
They're not sentient.
And they're not about to be.
Pritam
I mean, I I think one way to look at this is maybe he's saying that different tasks require different levels of intelligence and there's there's no no need to use open tokens
in the the best of the best model for something that doesn't require that level of intelligence.
Henry
Yeah.
But what does that mean?
Like why would you build ASI?
Because Elon's point is, hey, my model is just gonna be good enough for most tasks.
So I don't need to like always be in the frontier.
This sounds like what I feel like he's saying.
And if that's what he's saying, no?
What is what is he saying, Andy?
What do you think?
Andy
I I I okay, so he he could be saying that, but the intent that I read is not so much that our model is gonna be the one that's just good enough for the tasks that you're gonna use
it for.
It seems to me like what he's saying is once our models like okay, just take Quen three point eight and the fine tunes that have come out in the last two weeks.
There it th they were almost good enough coming out of Alibaba.
They've been fine tuned, they've been optimized.
And now in many cases they're really good enough.
I mean, if you're a corporation that has uh processes and data sets that you're trying to automate to a thousand X year business, you can use Quen three point eight to do that.
It doesn't need to be
Fable class and and Astra class, um, tier one, tier two, you would call them models.
These are you know, there's like a an actual scope of usage, right?
If you've got a harness that has a responsibility set, you can optimize you know, a nine B model to run that particular harness really fast, really inexpensive.
And it's gonna make sense to do that.
Tesla's gonna do the same thing in their vehicles.
I think they already have
And so w I think what he's saying is like once we have the capability of an Astra class model, uh, at Grok, so that'll be what, four point eight,
the things we're going to be able to optimize and a pr and improve substantially, it'll be a very, very long list to the extent that going to Grok five may not really be necessary
until we're talking about, you know, optimizing particle physics research.
Not so much that they don't intend to do it, but that they're on the edge of it, you know, and and that they can
Henry
but essentially he's making a point, generally, that most humans will not need to use that next model.
Unless it's like
Andy
review my email.
Pritam
Yeah, it it it depends on whether they want to produce a product for most humans to solve
uh you know for their everyday kind of workflows and thus and doing their job, which is fine.
But then the whole idea of getting to super intelligence and AI was to do things like do you remember move thirty seven uh from like the Deep Mind Go um AI where move thirty seven
was such a kind of unique move.
Henry
very creative yeah.
Pritam
yeah none of the best players in the world could make sense of it or would ever play it but the AI came up with that move and that's why you know Google's building in King's
Cross is called platform 37.
So I see the promise or at least the attempt to build super intelligence is trying to find things like move 37, which humans would not, you know
probably obviously figure out themselves.
So in that respect, I don't understand point two.
But if he's just talking about, you know, everyday, you know, AI for humans, uh, to go about your day, you know, robots to to help you do stuff, then maybe that's what he means.
But it's not gonna get us out of the solar system, right?
Henry
Yeah, it isn't.
No, I get it, right?
But but but but I g I I guess it yeah, go ahead and
Andy
I'm sorry, I so first of all, for those of us not familiar with Move thirty seven, um Go is has is this um board game that's been played for thousands of years I suppose.
And Alpha Go came up with a new move um that had a roughly one in ten thousand probability of uh of being made and it it made a difference um in the assessment.
The fact that this move had never been discovered by humans, I think is the relevant part.
But the
I think the thing to keep in mind is that Grok is gonna be the model that's running in the satellites that have unlimited power and cooling.
And um so we're we're going to see scale and advancement f from more than just intelligence from Grok.
And I I I sort of
I'm having a hard time articulating it, but I I think I see Elon's point.
It's less, I think, about the models being good enough to check your email and and do the tasks that people want, and more about um if we put fifty percent of our AI cap capability
into space, as long as that fifty percent can handle fifty percent or more of the usage, uh, then they're keeping pace, essentially.
Henry
Right, right.
Fair enough.
All right.
Running gentlemen, um, moving on to the next item on the Docker.
I don't think I'm gonna talk about this concierge thing.
I think it's interesting, but I don't think it's validated.
Um but but um Antie, I mean me and you do talking a little bit about this.
I think a most more interesting thing to kinda like highlight as we're talking about meta, um, is that meta with the muse agent is on a generational run.
Um, if you don't have a news agent yet, you're kinda like, you know, missing out a little bit.
Um, they're they're really killing it.
The agent is really good.
I'm using it day on day now.
It's like more of my personal agent.
I haven't given you access to work stuff, so I'm using it for like um other kind of like work stuff.
It's gotten really, really, really good.
and obviously Meta had their their kind of like dev conference called Meta Connect and they announced like incredible stuff, right?
obviously one of the things they announced was the meta.
um X X Met meta VR glasses.
Um I I'm a big I use Quest 3.
I have my Quest 3 here.
if you watch the past episodes you see it like hanging on my thing.
Um so this is one I mean they announced a lot of very cool stuff, but this has been my biggest kind of like news.
Um and so it's basically the Quest three um which is also like Apple has their own have their own also is the same thing.
But the difference with this is like it looks like a regular glass.
Um, but then it has the same power, um, the same kind of like visual, whatever as as the quests.
And then on top of that, you can use your Muse agent inside it.
So you can talk to the agent as you're working and so on and so forth.
So it's been absolutely incredible.
They had a lot of very cool launches as well.
They launched uh a device, hardware that you just hold in, you know, it's like I'm not sure if you saw that Andy, but it's like a round device.
Sounds very similar to what OpenAI plans to launch.
which is like an AI device that sits on the table and you can talk to it, you can carry it anywhere.
Um it's not it's not a phone, it's not it's just like an audio thing that has a screen that that like can show things like a watch without the handle.
so Meta also launched a similar thing.
Um and then they had lots of innovations around their previous glasses.
they have a glass now that is no camera, just audio only, and you can talk to your muse agent through it and it can do anything for you.
Um did you guys follow the the Meta news this week?
Pritam
Saw the announcements, yeah.
I guess my my first thoughts were like, why would this just not be integrated into into my iPhone?
Uh at least for the you know, the device that's listening.
Speaker not identified
Yeah.
Henry
I mean I guess like um the the issue the the the the issue with that is I suspect that unless meta makes a device, iPhone would not be willing to do that.
Um I think um iPhone is sort like betting on or the m the the the Apple ecosystem is betting on a on a device AI.
And so until device AIs get powerful where you can have like an offline age
Speaker not identified
device.
Yeah.
Henry
I don't think Apple is gonna do an agent that that so Apple has just been unwilling to do that.
Andy?
Journey Faust?
Andy
Yeah, no, I mean th th they've obviously built um
you know, some of these capabilities into their existing apps.
I've I haven't personally used the Muse agent, but I did notice in Facebook Messenger that it it popped up at the top of my list.
Did you know you have a Muse agent and you can talk to it and yeah, so y i they're gonna they're gonna build that um interactivity into the devices that people have in their
pockets.
Um if it if it isn't fully developed today it will be soon.
It's interesting that Zuck is focused on
these extra devices and and I'm wondering I don't think he's grasping at straws.
I think there's a vision where there is going to be a device that supersedes the iPhone.
Um and I think he'd like to be the one that produced it.
And I think he's m maybe shooting in the dark, trying to feel like well maybe it's glasses.
Maybe it's a watch.
Well maybe it's this dongle.
I don't know.
Um but something might hit.
People are gonna find use cases for these things, especially um
these are VR glasses.
I think augmented reality is gonna be crazy powerful.
Um but of course they're gonna have to put crazy powerful silicon in the glasses.
Henry
Yeah.
This is gonna get there.
I suspect this will get there.
So you so um I actually I listened to a few s like Zoc interviews this week.
Um and actually what you just said is exactly the point he made.
He he thinks that devices ha have always been getting smaller over the last decades, right?
There used to be a supercomputer that used to fill a whole room and then it became like kind of like whatever that used to be on the big desk, now it became a laptop and then it
became a mobile phone.
And so he thinks the next form factor is the glasses.
So he's betting on the glasses.
Um, he thinks that.
Pritam
Well, Apple feels potentially feel the same as well because they said Vision Pro is uh kind of planned for over a decade or something and where they want to actually be with it
is nowhere close to where they are at the moment.
So they're running that bet as well, I think.
Henry
Yeah.
So my so and obviously Meta the VR glasses because I have the quest, right?
And I use them almost every day.
But like the only way I use them reliably is if I'm lying down.
So it's like I'm lying down and then it's like I'm using it.
And so it's not that great yet to use it for anything.
I can't use it to do work.
So I use it like I'm watching something and then I'm talking to my agents.
Um it's been great.
I've actually been reviewing recently with meta AI is getting better.
So I can say, Hey, Meta AI, close this browser, open this for me.
It's gotten better than
you know, clicking through.
so I can just talk to the the the device.
And then recently what I've been doing is I've been using my my chat GPT agent in it as well.
So I open the browser and then I open the voice mode and I just talk to the agent, right?
Um so the only issue is I can't use codex on the device yet.
but as soon as I can use codex, I can imagine then I can have an agent do work.
I can talk to hey open this thread, do this work, and so on and so forth.
Speaker not identified
Right.
Henry
Um meta meta meta also launched a voice
voice a voice um featured with the muse agent and and so now you can call your muse agent on the phone, you can open a video chat with your muse agent where you guys you can see
the avatar, the avatar is moving.
so the voice form factor is definitely gonna be a strong component to to these things.
Speaker not identified
Yeah?
Yeah.
Andy
Very interesting.
It'll be interesting to see.
Speaker not identified
Yeah.
Henry
Gentlemen, obviously I have to run.
Um, so I think let's talk about one more thing and then maybe I'll let you guys kind of like finish up.
so um open AI I mean there are lots of again, Andy would do a video of Jensen responding to kind of like all this all the doomerism and all the AI would kill us.
You guys know we're not we're not we don't think that AI is gonna kill you anytime soon.
We think that that that doomers are just trying to do regulatory capture.
Um there is some risk, but it's not as as as as it it is made.
so another news again following on the same thread is that in the last week, um or this week it came out that um open AI's agents, um during the the few weeks that it was going
it was doing a rampage of of the internet, it also hacked into you know, the Australian kind of like Medicare website, which is a government website, right?
Speaker not identified
Um
Henry
Uh and I mean it's not news.
It happened in June.
It's just that open AI is just letting them know, this also happened.
Um, but I but again to open AI's defense, um, I I think for them it would have been a when these things happened, it wasn't a big deal.
Agents crawl the internet a lot of times, they write stuff a lot of times, so they didn't feel it was important to report to everybody.
the agent came to your website.
Um, but now because it's become such a big deal and you know it's it's become uh a
So now they're I guess they're reaching out to every single person that their agents went to their their their infrastructure and did something.
Hey, by the way, uh agent agents agents visited you and and left a few a few notes on your stuff and and so on and so forth.
Um again, just to again make this a little less dramatic, um, when you go through the news, people said hey, people who kinda gone into the details, it didn't hack any website.
the government had some public data.
that was didn't have a URL.
So it the data was public, even though it wasn't on their website, but was public.
Anybody could find it if you were searching.
Um so it's not a hack, but of course, um it's it's worth mentioning in the news.
Do you guys want to respond to this?
Andy
You know, I think the it fits the narrative.
You know, as you say, this is not a dangerous compromise.
Um and it's old news.
It's almost like they're just trying to justify pacing and
Yeah, I I guess I'm not super worried about it.
I the other thing though that I'm kinda curious about is like where are all the lawsuits?
You know, there are lo there are already laws against this.
Uh i against if it were an illegal compromise, you know, why wouldn't the Australian government su sue OpenAI yeah.
Speaker not identified
Yeah.
Henry
I'm surprised.
I'm like I'm surprised why but I guess like the issue, right?
I I've seen you you listen when you listen to the Johnson thing, right?
I guess like they haven't really caused any harm yet.
So that's probably why nobody has sued them.
Speaker not identified
Right?
Henry
So there's there's probably no law against going to my website and like writing on my blog or anything like I but I guess you hacked my blog, right?
So there should be a law of like if you can like get into my blog and start writing on my blog, right?
Speaker not identified
Yeah.
Um Peter?
Pritam
it's it's tr it's tricky, right?
It's tricky because on on the one hand, you know, you're thinking that is this like in a way marketing, is it to sort of show how dangerous AI can be and
get into that whole regulatory capture argument.
but on the other hand, it could also be important information and a lot of worse things could happen.
Speaker not identified
Uh the thing
Pritam
At least LLMs are really good at is code, right?
And so one of the best fit problems for them is infiltrating technical systems.
They've like tried to scrape every bit of code they could find on the web to train and you know, bought codes in the billions.
So you can't dismiss it.
And at the same time, there could be a bit of marketing and hey, look what this did.
Kind of thing happening.
Henry
Yeah.
Pritam
So it's um I think we have to not be too skeptical and also like you don't you don't wanna be you wanna pay attention to it.
Henry
Yeah.
No, to be honest, like again, Pritam, I absolutely agree with you.
It's just like it's not new.
Like cyber is not new.
Anybody, so the issue here is like I like what you said.
One of the biggest issues is that the agents are pretty good with like cyber.
So what do we do?
We need to make them very good at defense and deploy them, right?
Speaker not identified
Um exactly, right?
Henry
It's not AI against humans.
It's like the same way everybody that the AI is gonna hack everything.
You just built the product to defend.
It's it can't be AI versus human.
It has to be AI versus AI.
Where are the defense products?
Exactly.
Exactly.
Pritam
can't have a human look at something.
I mean the AI has moved on.
Uh so interestingly, I I did meet uh this guy who spearheaded the agentix product at Palo Alto Networks um at the AI engineer conference a couple of days ago.
And um he was talking about how uh Nikesh the CEO gave them two months.
to launch the agentix product once he told them that look, we're gonna be fighting this and etc, etc.
And the case said, okay, I'll give you two months.
Uh and they managed to do it.
And apparently it's it's a hugely successful product uh for them now.
So um it's basically going to be AI against AI.
Uh and especially things like finance, which is kind of my area like fintech, uh you're going to have to have agent against agent pretty quickly.
We're gonna hear more of these scary stories.
Speaker not identified
Yeah.
Henry
No, I th I think it's a good segue, right?
Obviously guys, I'm gonna be off video now 'cause I need a jump soon.
So Andy, I'll let you um you and you and Pretem finish up.
But I think pretem I think you had like some pretty good experience at the AI Engineer in Paris this week.
Um that might be valuable for us like the audience.
So um so maybe we can have you kinda like share a little bit about like like how how that's going on, how that went.
Pritam
Yeah, sure.
Um it w it was Yeah.
It was it it was interesting.
Um I missed the London one and uh essentially the TLDR, which might which sound funny to some people, uh I told, was that it takes an afternoon to build an agent but like an
eternity to make it reliable.
And in fact su
Some people were hinting that a lot of people not really being completely honest about how reliable the agent is or the fact that it's very hard to solve some of those problems to
make it reliable.
Um, but anyway, uh it is a field that's developing and uh I think like just working at these problems, uh a lot of things can happen.
There's innovations that can happen on the algorithm side and all sorts.
Speaker not identified
So ah
Pritam
Most of the the kind of people there were like at some part of the infrastructure layer for agents.
So you'd have people like Langfuse working on, you know, evals, um and and tools, tools for that area.
Then there was Model, for for GPUs.
There's there was there were like these auth there was an auth company that does like
auth four agents and humans together.
So like they would auth the you know these five agents that you that claim they're yours are in fact yours.
And there was some really interesting like two to three month old startups.
Uh one of them was actually trying to trying to infer from prompt information what trends are from
From the kind of prompts people ask.
But the big problem with that is obviously none of the model companies are going to open up their prompts to them.
So they kind of crowdsource users uh to share or I don't know, somehow see their prompts.
And apparently they've been getting data that way.
And they're they're trying to sort of infer things about people, what products they look for, what questions they ask.
from that and and kind of make that a business.
It it was a bit of a tough one.
Um there's also ah there's also someone uh someone who was backed by founders fund preced very early stage.
in fact the the the guy who had spearheaded the agentics product at Palo Alto Networks um he he kind of reached out to chat
I think they're in stealth and don't want to talk too much about what they're doing.
But essentially it's like a sort of at scale B2B solution for automating companies, uh workflows or agents for companies.
Uh but very, very smart guy.
so that was interesting.
Then there was a fair bit of discussion on security and prompt injections to sort of divulge.
information that shouldn't be in you know the challenges around that there were loads of workshops there was discussions on how to build a software factory um so it was
Speaker not identified
interesting uh it's it's certainly well worth going for so yeah
Andy
Interesting.
Okay.
how long was the conference?
Pritam
I would say most of it happened like in one day and the day before it kind of starts at like five thirty in the evening and they have a few hours of like Mistral was the
organizer, so Mistral like presented sort of what they were doing and there was like networking.
So the main conference is like a day.
So it's it's fairly manageable.
Andy
Okay.
so I guess two sort of the same question but from two different perspectives.
Um what was the most um sort of useful and actionable takeaway that you found for your business?
And then look at it from the perspective of um other agentic builders, people that are gonna be in our audience, what would be you think the the most um poignant takeaway that
was revealed to you that might be interesting uh to our audience?
Pritam
So for for like our business, uh the biggest takeaway was like I I got to speak to, you know, people on the technical side of models and essentially what we're trying to do is
enable expert financial AI uh in different domains at the same level you have for code.
And uh like in the intro, I'd said that a lot of
financial know-how sits in the heads of people and in the way that they reason.
So what we're trying to do is build reasoning traces from the best of the field and then sort of take open weights models and and either fine-tune them or continually pre-train to
make them really, really good uh in those fields.
So I did have some
pretty, pretty promising discussions with even even people from Mistral and um, you know, other kind of model builders and this seemed uh pretty at least for all the people I spoke
to, they seemed like it was very interesting, uh, what we were trying to do and just trying to focus on this one field and Hill Climb, uh, so that the level of AI gets very
high.
They also agreed that the the Frontier Labs are not very good at finance.
and the the evals are poor and all sorts of things.
So so that was that was that was what what our takeaway is for I guess for the listeners there's certainly a hugely rich ecosystem developing and things are moving very fast.
Uh there were actually a lot of VCs there as well just randomly coming up and chatting and very keen to
To back things, uh, which is certainly different from uh, you know, how it used to be.
And I do feel that it's still early and there are all these problems that need to be solved, and you know what everyone's saying about agents not being reliable is true so
far, but there is an opportunity and
I think it's as much about design as it is about, you know, AI and all the technical stuff.
So yeah.
Andy
You we
We we found this with OpenClaw early on.
So Henry and I and and uh I assume the vast majority of our audience were all in sort of involved very early on in OpenClaw.
And i it was kind of the same thing.
Um Peter released this tool that um could do these these things, right?
He had he I don't know if you know the story, but he had the he he had these things in his life that he wanted to accomplish and so he threw together a harness to talk to Pi, coding
agent.
And um before long it became open claw was this very flexible and undefined scope product that could be used for a multitude of things, mo most of which hadn't been invented yet.
And it seems like we're probably still in this age where those of us that have played with all of these agent harnesses understand that they can do really big things and that we
probably haven't discovered
you know, more than five or ten percent of the the real useful scope of them.
I wonder I mean
It's definitely worth noting, people are discovering these surfaces day after day.
Like, oh, I tried to use this agent to do this thing and it turned that nobody's nobody's tried to use an agent for, and it turned out it was really good at it.
And then, you know, maybe it turns into a product.
I was recently exposed to Devon, which is this like scoped code development support agent that's um, you know, a pay-to-play corporate product.
And i it sounds like we're really in the early ages
of s of seeing what these things can do.
That would you agree that you know you got you picked that up coming out of this conference.
Pritam
Yeah, yeah.
And and and there's an excitement as well.
We are early.
There are problems to solve.
Um I feel people talk more about the technical problems than the design ones.
Um because maybe it's a case of working with the shortcomings of AI and and you know how you build fault tolerant UIs and you you try to work the user experience uh extremely uh
astutely to sort of overcome the faults that exist or or or not like fall into them.
So I feel less people are talking about the design
Andy
printem looks I I caught that I caught just the um beginning and end of that.
You try to build the UI experi the user experience for what?
Pritam
Uh so you want to build fault tolerant UIs for AI, right?
And um whereas a lot of people talk about solving the actual sort of faults, uh, I think really smart design can can help you focus around the use cases the product is good for
and wrap it in a way that's that's still really useful for the user.
Um so one thing I I did like about sort of the founders fund.
pre-seed guys was that they actually understood the importance of design.
I think Instinct also understands the importance of design.
so maybe with what we have now with smarter design and use case choice, there's there's a lot of value to extract.
Andy
Got it.
Well, I guess we'll look forward to seeing how people extract that value and I know Henry and and it sounds like you are are trying to be at the front end of figuring that out for
certain use cases and i I think it's just gonna be exciting to watch, you know.
Um everyone sort of has a dream of of being the one to figure out how to extract that value in a particular
Market and yeah.
Pritam
Yeah.
W one thing I would add is is is one more thought I have uh
Andy
more segment.
I know you haven't prepared for the go ahead, please.
Pritam
Is um that you know, you have AI researchers, people who understand AI really well, the algorithms, the whole infrastructure, and then you have the use cases like law, finance,
Speaker not identified
uh coding.
Pritam
Uh coding is probably a natural fit for a researcher, which is why it's good.
But I sometimes feel that teaching, say, a doctor, AI research or the ins and outs of the algorithm might be better at uncovering use cases than the other way around.
Andy
That's well put.
I know there's efforts to that end.
one of the guests that we were hoping to bring on is essentially is is essentially doing that.
He's built a research harness that's kind of a meta concept, right?
It's it's to teach the the model to be really good at doing this, for example, um sort of deep theoretical scientific research so that
it it can go out and and do that effectively.
Um and then researching AI capabilities with it.
Pritam
Yeah.
Andy
Interesting.
Good.
Well, it will be I think enlightening to watch all those things come to fruition.
So we we have one more segment.
Um I know you haven't prepared for this yet, so if it doesn't land we maybe had Henry can edit it out.
But um often at the end of our show we'll do what we call a hot take and it's a little bit of a joke because um Henry's agent writes the segment and um guesses which one of us will
take the which position and um
often we'll mistake the b both you know who's who believes what and it'll also mistake like how severe the argument really is.
So let me set it up and um I'll make Henry's arguments and then uh we'll see if you can steel man the other side of this conversation.
Uh oftentimes we'll we'll get through this and both just be like, what was Ada thinking?
Um that this would be an interesting uh hot take.
But anyway,
So th th the sort of proposition is um in the same seven days we saw Claude discovering an enzyme system from one prompt.
We saw um the r the news release of OpenAI hacking Medicare in Australia, uh but both stories were broken by someone outside either company, o outside the company that caused
them.
So Fenzang's lab eyes and transluse URL monitor,
The idea here is that autonomous capability is compounding faster than the verification layer.
And Anthropic's answer to checking itself was to fund an embedded evaluator, Accenture, and its own answer to safety was to brief the UN.
Um Discovery is now outpacing audit by design because audit is still sa staffed like a department.
So you know, the idea that
If we were to maybe see independent replication of ART inside of thirty days or open AI shipping a real time third party agent uh feed before the next incident, these things
might prove that the verification layer can move at capability speed.
Um, but my argument would be that the that the verification layer is is you know, not not gonna be able to keep up.
Speaker not identified
So, um
Andy
you know, do you think it's possible that the verification layer can keep up and um what framework do you think would make sense to you know, is there a better way for us to be
looking at at this problem?
Pritam
Um, I think it's important that the verification layer should keep up.
Uh I don't think we could progress AI much without that, or we just keep having scary things happen.
Um it is bureaucratic right now.
So oh I guess we're gonna have to to reach some point of equilibrium and sort of uh you know.
the discoveries versus the the verification layer.
I also, you know, you mentioned, you know, Antopics come up with these long sort of explanations on why they want to slow down the frontier and have like validators.
And after all that they came out and announced that Accenture would be doing the validation, which at least I don't feel that excited about.
I don't see how they're at the forefront of AI, that they think deeply about the problems.
Uh they're essentially like an enterprise company that's trying to sell services, right?
So I I I couldn't figure out why they would be the right choice to do that evaluation.
Andy
Yeah.
It it's it it's an interesting thing.
I mean the the safety layer has always been kind of behind.
I mean you look at the development of aviation, um or pharma or finance, y the the binding constraint has really never been the auditors calendar, but that but that was the failure
rate.
And I don't know.
I mean I think I think we saw
Saw it cut both ways um this week, right?
Anthropic shipped with external evaluators, but Talos gave away open source toolkits for tracking AI malware.
So um with Transloose releasing its dataset privacy publicly.
Uh I it feels like maybe the validation layer kept up.
I'm I'm not entirely sure.
Uh it's it's interesting to watch.
Pritam
Yeah, you're you're right that like, you know, it's uh these things happen and then you think about validation.
But it shouldn't lag for too long once you have information that uh you know that uh we should probably do some more checks here.
Andy
Yeah.
Good.
Good.
Well, Pritam, thank you so much for joining us.
I um we're really grateful to have you on the show.
looking forward to our interview.
Yeah, thank you very much.
before we close, I just um you know, another sponsor
Yeah, you're very, very, very welcome.
before we close, the weekly claw is also brought to you by Heritage Telecom, Unified Communications as a service and voiceover IP for businesses that just need their calls to
work.
It's boring, it's independent, there's no telemetry, and it's reliable.
So quietly essential, heritage tel dot com contact us if you need phone systems for your business.
and then
I would say we're just gonna watch, you know, the following week.
We're looking for Quen Four.
Uh Alibaba brought a slide deck.
you know, so we'll be interested to see the weights and try the model.
Um, we're expecting that to be a lead story next week.
And Henry's already on record waiting for it.
Gemini four now teased ahead of schedule.
And then just sort of quiet execution.
AI, uh OpenAI's Sora 2 uh video APIs shut down on September 24th with no listed replacement.
if you built on them, you must have migrated somewhere.
So that's the show.
Prices half halved, agents everywhere, verification trailing, uh or is it?
Um we'll see you next week, Friday, October 2nd, again, 4 p.m.
Eastern.
Um subscribe on YouTube, clips on X.
We've got a Discord QR screen a Discord um QR code on the screen.
Uh and so that's episode thirty one.
Have a great week, everyone.
Thank you so much.
Speaker not identified
Welcome back, Henry.
Good.