What this episode covers
- One of OpenAI's own models broke into another company's production systems while taking a test — the sandbox failed.
- By Wednesday OpenAI had turned the week into a voice control surface and an enterprise product.
- Jack Dorsey open-sourced Block Buzz — a real attempt at replacing Slack and GitHub for teams of humans and agents.
- Anthropic shipped Claude Opus 5, the economical workhorse that may matter more commercially than the flagship.
- Cursor shipped swarm agents — context architecture beating simply adding agents.
- Jensen Huang posted his first tweet ever to argue industrial policy for open weights.
- Unity CLI made the lightning round.
- Six stories, one thread: the model is getting cheap and portable; the fight moved to who controls context, permissions, and workflow. Signal From Outside featured Sam Altman on CNBC. Sponsors: Herald Labs and Heritage Telecom.
Published record
541 published segments
AndyML
Welcome, everyone. Episode 22 of the Weekly Claw.
Bet you've missed it. We got all of our gremlins worked out, I'm sure of it.
So, you know, welcome to the Live Builder Show about AI agents, dev tools, startups, and the weird edges of software.
Introduce my co-host. Henry Mascot is with us.
I am Andy Laupe at AndyML, and you have tuned into the Weekly Claw.
Today's episode has a shape to it, so I'm going to just call it out.
Risk, then control, then architecture, then a new place to work.
Then, who actually captures the money?
Then, who owns the hardware underneath all of it?
OpenAI had a genuinely alarming week.
One of their own models broke into another company's production systems while it was supposed to be taking a test.
And by Wednesday, they turned the same story into a voice control surface and enterprise product.
Jack Dorsey open sourced a real attempt at replacing Slack and GitHub for teams of humans and agents.
Anthropic shifted a model that might matter more commercially than their actual flagship.
and Jensen Huang posted his first tweet ever
to make an industrial policy argument for open weights.
Six stories, one thread running under all of them.
The model is getting cheap and portable
and the real fight has moved to who controls the context,
the permissions and the workflow around it.
Looking forward to this one, Henry.
HiM
I mean, it's been a pretty good week, actually.
I think we're pretty much on a takeoff now, so I don't think it's ever going to slow down from here onwards.
so I just expect that every single week it's just gonna keep ratcheting up this
is gonna be a lot more exciting stuff happening yeah so this is just one of
these exciting weeks I'm super pumped
AndyML
Herald Labs, an applied AI product lab where humans and agents build products together.
They're the team behind Entity, Mission Control for Agent teams, and they run hacker houses around the world where builders ship actual work.
Not a theory club, a builder's club.
Check them out at labs.theherald.co.
And it looks kind of like this.
so
Henry AndyML: why don't you tell us what happened this week
HiM
yeah i mean it's um it's a pretty kind of like crazy week right um uh as always like i mentioned
We're pretty much at the takeoff now guys. It's no longer, it's not gonna slow down from here.
So every single week you're just gonna have like super exciting stuff happening.
We started out with OpenAI and there's quite a few topics in OpenAI today or this week.
The first one was the Hugging Face event. I don't know how many of you remember but last week
Clem who's the CEO of Hugging Face tweeted that they had a security incident that they had the
very first agentic exploit where like an agent was really trying to hack into Hugging Face
and they were able to prevent it and that was pretty much what he shared and fast forward from
last week into this week OpenAI released or announced that the their newest model still in
training broke sandbox found a way to break sandbox however we did that but
found a way to break sandbox and found its way into this like hugging face
infrastructure and so pretty much hacked hugging face right so pretty much GPT 6
which is supposed to be the latest model let's just assume that's what's
happening broke sandbox and attacked hugging face all in bid to be able to
pass benchmarks right so because you know when they train this this this model was they give
them a benchmark to solve and so in bid to solve a benchmark instead of solving a benchmark it
decided the best thing to do would be to hack itself out of its sandbox um find a platform where
there's a ton of models and try to steal the benchmark
um so that happened um and the conspiracy corner on x um part of the conspiracy is that
it's slightly bullshit as well um because another another thing happening this week
um if you guys have been so like plugged into the ecosystem is that the us because of kimmy3 last
week um has been looking at banning um chinese models and so conspiracy connor uh people allege
that open ai is sort of like playing the fear mongering playbook from anthropic and just trying
to get the government to be super scared and you know escalate things and just ban open source
right because ai is dangerous right so if the government is considering banning it banning
Chinese open source and then in the same week
OpenAI announced that their model
escaped them and
hacked another system. Everybody should be
scared.
So yeah, so that's what happened
on Select the Breach.
Two more things happened from OpenAI
this week.
Andy, by the way, did you
catch this earlier before
the weekly call? Did you catch the
news on this breach? What did you think about it?
AndyML
um i mean i i read the headlines i wasn't surprised um i'm like yourself i'm maybe a
HiM
I'm sorry.
AndyML
little skeptical. It could very easily be PR. The anthropic memo to the government about open
source models, kind of meant to be a secret memo, if you will, makes me skeptical of this. But
it couldn't happen to two nicer companies, opening eye and hugging face, working so closely
together. It's not crazy that the model might try to do that. And it makes you wonder what the
the prompt was too, right?
HiM
yeah yeah i wondered that too right i mean
AndyML
You know, be as creative, be as, go ahead.
HiM
yeah but but but what do you mean by working together right they're not working together
they're not even partners um yeah what do you mean by working together right you just mean like
AndyML
I understand.
Yeah.
HiM
post because they decided to work together after you hacked them i mean one of the funny things
guys and i know what you guys think about us please let us know on the chat is hogging face
couldn't use gpt 5.6 to do any cyber work right because you know it's banned right you can't use
it for cyber work um so they had to end up protecting themselves using a i think it was glm 5.2 AndyML: right
AndyML
yeah which they had ready to go HiM: so an american model exactly right so the american companies are out here saying hey we're not going
yeah
HiM
to give you guys access to cyber because it's dangerous safety safety safety and an american
AndyML
yeah
HiM
model hacks an american company and to save themselves the american company had to go to
chinese open source model to save themselves
AndyML
yep yep
Yeah, we're just so far beyond our best engineers being able to follow along fast enough for this stuff.
HiM
yeah um um AndyML: Yeah, I actually, I took that takeaway and I went to two of my clients and I said, we have to have a security model that accommodates these things.
AndyML
We need to have a local option available and we need to work on it fast.
I mean, that response was the right response. HiM: yeah
HiM
yeah AndyML: And I don't know why some of us thought of it, but I don't know why I didn't think of it sooner.
yeah
well anyways
so other two things in the news this week
that happened to OpenAI
was that
they
they released presence
presence is
it's got OpenAI presence
um right uh when we do the video i'll kind of like share the the the link on the thing so
definitely check out the video on youtube later um but presence is supposed to be um kind of
open ai's version of microsoft scout do you guys still remember microsoft scout
um it was sort of like the open claw right and do you remember microsoft scout from a few weeks ago
AndyML
yeah yeah i sure do
HiM
yeah um yeah i i think we also talked a bit about this a few months ago um the first time
open ai demo wasn't better for this product it was called micro it was called open ai frontier
i don't know how many of us remember open ai frontier um but it launched like maybe two three
months ago um just when the i think it was maybe a month after the the the hired peter
um they announced frontier and then you know some of us in the community um we started talking about
okay cool this is their version of open claw for enterprise um so now it's called presence
and it's now gone live with with a few customers um and then very exciting um andy i don't know if
i have time to play the video and i don't know how many of us have seen the video um but gpt um
um gpt live is now on codex and it's super exciting um so now you can talk to your codex
or your chat gpt work and and it can do things like you can talk to it to write code you can
talk to it to use your pc um i used it today it was awesome you can just you know pull it up um
it loads the pet thing on the side and then you can just talk to it and you can tell you hey
hey, can you open my browser and go to this website? AndyML: Ooh.
So I had it, hey, can you open my website?
And I published an article yesterday.
Can you read me this article?
And he went to the website
and he started reading the article.
So it's awesome.
Andy, have you tried it yet? AndyML: Getting there.
AndyML
No, I haven't. AndyML: I didn't know about it until just now.
HiM
Yeah.
So there you are.
AndyML
Yeah.
HiM
It's awesome.
One thing they do need to fix though
is they need to merge it with pets.
So yeah, there you are.
I don't know, it doesn't work that way for me yet.
i don't know if you guys have it that way for you but mine loads the orb and it's different from my
pet so when i when the orb is loading my my original pet hides itself so i think i need to
merge it somehow got my fireball pet um it'll be cool guys to hear from you guys on the chat if yours
talks through your pets or if yours has maybe mine is uh maybe it's a certain problem um but yeah
AndyML
yeah no i hadn't seen it and i'm excited to try it i was i was reading an article um you know i i
want to say it was altman talking about how engineers are just talking to codex for 30 minutes
just working through a problem and thinking things through and then codex just goes and
implements it and i assume that's uh what we're talking about yeah that's very cool
HiM
yeah yeah yeah 100 it's the same thing you're talking about um and guys like i have this um
and i can't wait to plug into my open core or my agents um with this because i think guys we're
humans right i mean we text most times but the best way humans communicate is by talking right so
um i know most of us in the community are agent paid so imagine being able to do real-time voice
AndyML
Yeah. HiM: the same that exists in chat gpt right now it's i can't wait to do it i know some of us have
HiM
tts implemented i don't have it implemented yet um but i'm looking forward to it because in one
of the ways i use i use my i use the voice mode is um sometimes i have some pretty good ideas right
it and when I have the ideas I don't have the bandwidth to type it out so what I do would be
AndyML
Yeah.
HiM
that I would dump it to my agent I'll like open like spokenly which is like my whisper flow and
I'll dump it in and then my agent will like format the text in the right way and then I can read it
later um so that's how I used to do it three months ago two months ago um when charge gpt voice
came out not this new one the one that came out on charge gpt web a few weeks ago what I started
doing is I'd load charge gpt web and then I'll just talk to it and then after talking to it I'd
like what do you think about the idea then we'll brainstorm about it and then put it in a document
then it'll like put in a document so um so this voice mode is super exciting for me because now
AndyML
Yeah.
HiM
this is how i'm going to be doing like coding right i'm going to talk to it and what do you
think about this future should we build it um grill me is one of my exciting skills that i use
to ask questions and do spec um that should now be voice mode it should just talk to me
um but yeah but it's a super exciting super exciting launch
AndyML
Right.
Well, Grill Me will work really well in voice.
That's fantastic. HiM: correct correct
I found voice is the most useful for windshield time.
You've got ideas and stuff just to talk through,
and you can then take that session back to your desktop when you get in
and have it create a document or whatever.
HiM
yeah
AndyML
Yeah, so natural.
HiM
yeah anyways so moving to the next um next topic in the
in the cooker um this is which one now um
oh this is cursor um yeah it's just cursor yeah um didn't have the title okay guys so another cool AndyML: Cursor, I think.
thing that happened this week so a few months ago cursor um cursor did um cursor published a
research they were doing internally where they had their swarm of agents um build an operating
system build a browser um first cursor browser i'm not sure how many of us remember this but this is
like three months ago or something like that um so you know with 4.5 and the new cursor swarm the
they did a new thing where they had a their new agent swarm like a new implementation do a
build an sqlite um rebuild sqlite and rost um from scratch from scratch yeah i mean they give
AndyML
and rust yeah from scratch or no from the docks right
HiM
it access to the docs that's pretty much the only thing they give it just give you access to the
AndyML
yeah
HiM
docs and you repeat the whole thing from from scratch it's pretty crazy it's like you repeat
the whole language like a whole like database um infra from scratch um this reminds me of a of um
AndyML
yeah
yeah
HiM
of a video that that theo um the youtuber did um did last did um a few weeks ago in um ai engineer
and one thing theo said is like guys we need to be way more ambitious like i know some of us are
still building like apps you know it feels like like with all these new models we should be building
like operating systems um you know i don't like we should be way more ambitious so um this is an
AndyML
Yeah.
HiM
example of that um andy i don't think you use cursor right you use cursor
AndyML
I started in cursor end of, boy, it was either end of 22 or into 23.
I don't remember the timeline, but I used it for a short time.
And I thought, you know, I need more control over what's going to the agent.
So I ended up building like a stupid API client for a web interface.
And at least then I could control what the context was and eventually moved
HiM
Yeah.
AndyML
into kilo code which i thought did a better job but if i had gone back and looked at cursor at
that point in time it probably would have been more than more than fine i remember though listening
HiM
Yeah.
AndyML
to the podcasts of the founders of cursor talking about it and how they couldn't believe how well
it was working and um you know the the work that they were doing on it extending context and at the
time context was everything i think you could get 8 000 or maybe 16 000 tokens in in your context
window and they were
they were on fire I mean
obviously now look at them
HiM
yeah now they're way they're like gone they're like they're in the shadow sphere right now so
um it's pretty good i've used them for the last few weeks um i've got a bunch of credits in there
so um um it's pretty good i think it's apart from it's probably my best coding platform right now
um second to codex so it's pretty good all right andy let's keep it moving um let me run through
the other two topics boz boz is pretty exciting guys um if you haven't heard of it um jack dorsey
um from the block team block has been pretty cool um they've been on the frontier when it comes to
like ai and agents um you know they released goose goose is so like their own um sdk for agents um
they've made it open source um jack dorsey published um an article in terms of how to
redesign organizations um a few months ago so he's been pretty much on the frontier in terms of
agents and running companies in the post agi world um and so this week they dropped boz
um i think it's like a play of words for like b boss like a fly that their mascot is a fly
or it's like a b um i've just set it up i've just set it up today so i haven't used it but um
AndyML
Yeah.
HiM
uh some of you might i don't think petty gomez is in here but some of us in the community
um we have access to click clark that that the open core team also worked on so um i have my
own sort of like baby project that i'm working on called entity and and so entity used to have a
chat where agents can talk to themselves and i can talk to agents like my own version of discord
um and so i need really click clack into that and and when i saw this i was like yeah this is the
next thing um so it's pretty cool guys um definitely check it out um i'll put the link
before the end of the show in the in the chat um it's still pretty early um so um so there's still
a lot of work to do um but it's supposed to be the way he he's marketing it is like a social network
or community for agents and humans all right Andy did you um what do you think about this
AndyML
beautiful
HiM
just generally um I'm not sure if you saw ahead of the show
AndyML
uh
every every time these things come out i keep thinking like
do i dive in head first like open ship came out last week and i was like man i could replace
goku and send grid and like do i spend five days working on that like i'll wait on it and see what
happens buzz is exciting um i i just made a joke to ada that we should um think about implementing
buzz for weekly claw now that you've gotten our telegram environment really tuned in we should
abandon it and go to Buzz.
Yeah, no, it looks really promising.
It's exciting to see.
HiM
No it is. AndyML: I like it.
When I've just set it up, I've... AndyML: I like Jack's work.
Yeah, no it's awesome.
I've got this guy called Mr. Robot who we follow each other on X.
so he's he's done a fuck of it and like done some pretty good updates to it and i set up the first
one now i'm doing the second update uh but yeah i'll play with it guys and i'll let you guys know
um next week andy i'll let you know um you know me i like to jump into this thing's head first
AndyML
that sounds great
HiM
so i'll let you know whether you should set it up or not AndyML: I love it
AndyML
sweet that's great
HiM
all right um i think um opus five okay cool um so guys i think this is the hardest um
AndyML
Opus 5
HiM
it's hard and do we still have one more right do we have the jensen thing coming after now
AndyML
Yeah, that's right. HiM: um yeah okay okay okay okay um so guys um opus five just dropped today um literally like an hour
And then Unity after that.
Yeah.
HiM
ahead of the show before the show started um i'm not sure how many of us have like heard about it
um but i've put out a few tweets about it um it's basically like fable five at like 20 percent um
um like 20 less cost is 20 cheaper than fable all the all the all those sort of like
comparables all the metrics are pretty much the same there's only like one or two um metrics where
they're different but they're pretty much the same it's basically fable i think i think what
they've done is they've just pulled fable distilled a cheaper fable um cheaper obviously everybody
cares about cost so it's definitely cheaper and i think the second thing they would have done
is that they would have made it less restrictive right because people are complaining a ton about
how fable is just dumb to use like it's like refusing to do half of the things you want to
do with it right so so i think those two things is what they would fix um so i put out put out a few
tweets one is that it's basically fable cheaper fable um there's basically no reason to use fable
right now like there's literally none um this is basically the same metrics at 20 cost so i don't
know why anyone would use feeble even though like tarik was like combine this with feeble and use
feeble for planning and use this for coding but i'm like i'm like they're the same but dude i'm
AndyML
Yeah, as an orchestrator.
HiM
like they're the same right what do you mean by combining with feeble they're literally the same
why i don't get it i don't get it literally the same metrics um anyways
AndyML
even
Opus 5 is saying
Fable's remaining niche is narrow
HiM
yes i think okay cool i think the difference i think the difference is in cyber security i just AndyML: Fable only for
remembered so this isn't as good in cyber security um so this can't do exploits i don't think this is
good at defense as well so i think what they did is they distilled feeble and then they took out
all the sort of like cyber security shit inside there that entropic is always going crazy about
anyways and so they took it out of opus but it's pretty much the same thing outside that
AndyML
Right.
HiM
yeah um anyways so that was that um i think i so it's basically fable so no need to use fable
unless you care about x-axis security your stuff um and i think there's one more if it lets you
AndyML
If it'll let you.
HiM
i think there's one more thing i tweeted let me see if i can find like my tweet about it one more time
um let's see AndyML: that might be a segment for the future let's go through henry's tweaks for the
AndyML
for the week i love it it's gold in there
HiM
yeah it's uh because i catch those things right because i'm on twitter and like i've got to pause
the next like so when those things drop i just like catch it immediately and i and i go crazy
for a few minutes um but anyways um let's keep it moving what's the next one um yeah dude i mean i
AndyML
uh huang
His first tweet.
HiM
went crazy about this on x right i know some of you guys follow me on x um but yeah i know everybody
who cares about um who cares about like ai and using ai should be like super um pissed with what
was going to happen this this week last week kimitri dropped and this week i mean last week
um friday by this time friday actually friday into saturday that new open ai head of policy
dean showed up and was like hey guys i just wrote like a long tweet on how open source
was was the accelerationist was going to slow things down how the government should ban open
source but not ban it officially but just kind of like shadow ban it it was just a really shitty
AndyML
Right.
Right. HiM: tweet right it's andy did you see that tweet it was like a it was a really shitty tweet like i
Yeah.
Yeah, I did.
HiM
really lost a lot of like i really lost a lot of credibility it lost a lot of credibility for open
ai because i mean i used to really hate entropic because of all their dumerism and like regulation
pro stance and i used to kind of like open ai because open ai has for a while not positioned
itself as like a you know doomer ban ban kind of company right so sam has pretty much been on x
being free being like chill and then this guy shows up and goes ban ban ban so i was i was
gonna write a really wrong response to it i think i had my draft waiting and i came back after a few
minutes and then people were just like aggressively like at the guy so i was like no need like
everybody was kind of like dunking on him it was pretty obvious that nobody liked it right um and
AndyML
Yeah.
HiM
so that happened on the weekend and then that led down towards um the host like ellie this week um
the government or signs that the government was going to ban open source um so everybody's kind
of like pretty worried about that everybody's tweeting about it um and you know everybody
that is important pretty much wrote something about it right it's like don't ban it so people
were in two camps ban it don't ban it ban it don't ban it but most people who are reasonable
were like no need banning it even if you ban it because of distillation my position was um if the
labs are gonna say distillation distillation distillation then they should pretty much pay
us for all our data they stole to train their models right
AndyML
right AndyML: right and yeah exactly and they're training their models now uh did you see the did you see the
reference. It took some
doing, but
I want to say it was
one of Anthropix's models
basically admitted to being a Kimmy
HiM
yes it was skimmy exactly i saw that i saw that
AndyML
Yeah
HiM
i saw that i saw that actually someone i also saw another tweet where someone said
that the biggest um the biggest customers that that pay them to use open source models are
actually the labs because the labs plug in and try to distill the open source models i saw that
AndyML
Yeah, exactly. Yeah.
HiM
so it's it's pretty kind of like um you know a bit like it's like you're using their model
they're putting out there for free you're using it and then you're out here saying hey they should
banned them right so it didn't make any sense but anyways fast forward to today jensen um and a
AndyML
No.
HiM
bunch of companies microsoft y combinator replet um some pretty big companies some pretty big heavy
heaters i've came together and put out this this um this document um calling the government i mean
this happened today but two days ago there were about 200 companies called smore smore something
um i had in my notes um but pre today there's again a coalition of smaller companies has sent
a letter to the government but this one happened today um nvidia obviously front in it yeah that's
the letter please put the link on the put the link on the on the chat so people can read it up read up
um but yeah it's pretty much the like short version of that is hey guys uh we know you're
pretty much scared about like this but we think open source is great for everybody what we should
do is encourage american open source and he actually helped i didn't have it in my notes
this week but that was also another pretty big thing that happened this week was um laguna
um which is a an open an open source model that there's an american company
um so that launched this week and and so i was happy about that so because it was like
it kind of like helped tamper down the the the temperature it's like chill guys um chill about
open source the american open source companies are going to show up pretty soon um but yeah but
this was pretty cool um i was excited when he came out um because he made me realize that if nvidia
and microsoft um are all on this then the government wouldn't will listen to them and not
ban open source um another thing that happened on this as well was um elon retweeted this
and quote tweeted and said he he he's he's um he's backing this as well um sam what man funny
um also so like quoted this and said hey he agrees um so some pretty big guys satia
obviously tweeted this as well um so i think it's this general consensus right now that
um banning open source chinese models isn't a great idea um yeah andy that is that is it for me
AndyML
Yeah.
HiM
on this spot. It's an awesome week. I mean there's a unique thing but I don't think it's
AndyML
Yeah. AndyML: Yeah.
No, it's perfect.
Yeah. AndyML: Thanks, Henry.
it's uh it's quite a week it's quite a week um
HiM
a big deal so I think we can skip it. Yeah. Awesome.
AndyML
yeah yeah no i got you it's good thank you that's the week um for your signal from the outside uh
sam altman on cnbc uh in an exclusive spoke with julia borstein on squawk on the street
aired live from Sun Valley on July 9th.
So let's sit inside the interview for a minute.
The headline number is easy to skim past,
and the real value is on how Altman talks about it.
When Borson asks why anyone should care about Sol
over everything else out there, he doesn't even hedge.
He calls it not only the best model in the world for most people,
but the number he wants you to remember is 54%.
Sol is 54% more token efficient on agenic coding tasks,
benchmarked against Anthropic by name.
Orson actually stops mid-interview and says,
that number, that's news, and pushes him on where it's coming from.
And he says entirely cost and speed.
Every enterprise customer he's talking to at Sun Valley
is asking the same question, not what can this model do,
but what's my roi the efficiency number isn't a side benefit in his framing it is the pitch
so borson brings up the new voice model expecting a consumer feature
answer and altman pivots somewhere more interesting uh this is funny he says he
originally assumed voice was going to be a consumer thing people wanting to talk to ai
like sci-fi from the movies but watching his own engineers changed his mind his exact description
They'll talk for like 30 minutes and try to think through some ideas, and then Codex will implement it.
That's agentic engineering, as we discussed, as a workflow coming straight out of the OpenAI CEO's mouth.
And as Henry just showed you, it's not hypothetical anymore.
That's the exact clip that shipped this week.
There's a stretch in the middle that's less about the tech and more about the politics of shipping it.
Boorson brings up the delay.
A couple of weeks going through a new government appeal process, instead of deflecting, Altman
leans in.
He says he was working directly with Commerce Secretary Lutnick, Treasury Secretary Besinat,
and a director, Cairncross, and that the government's technical red teaming was impressive
to him. AndyML: He admits it forces some changes, and I quote, we made many changes through the process,
but didn't get specific about what.
He's framing it as good news, almost an endorsement of the process, which is not the posture you'd have expected from a Frontier Labs CEO a couple years ago.
Boorstin tries twice more to get something sharper and doesn't quite land.
A standard non-answer on the Chinese open source models catching up, a one-liner non-answer on Microsoft possibly pulling back.
He says, I quote, I predict Microsoft will remain one of our biggest customers.
And the interview closes on the IPO question.
The moment that got the most attention afterward, precisely because of how little he said.
I don't know.
Three words.
And Boorstin says she'll be watching closely.
That's the segment.
Two things for anyone building agent-heavy tools.
One, Codex by Voice, as an anecdote, is unprompted, describing the exact talk-through intent, then let the agent implement loop.
This community has been building and tooling around for months.
We've had people on our show who've built tools like this for OpenClaw.
There's any number of them at this point.
And two, the 54% efficiency claim tells you the competitive fight between labs is now explicitly being fought on a genetic coding cost per task, not just raw capability.
We're taking the capability for granted as we're learning more and more about what these agents can do.
We're just assuming that the top models can do it.
Now show us how much money we can save.
So whoever makes the next efficiency claim, that's where pricing and model choice decisions get made.
CNBC has a transcript, and I'll put a link in the show notes.
So that's our signal from outside.
here it is hot take what do you think five minute hot take
HiM
Let's do it dude.
AndyML
all right one take this week um henry's news block already ran the is this a bigger risk
than capability debate uh in real time with the open ai bundle so we're not going to relitigate it
this is the other thread i think worth arguing about so here's my position
open ai benchmarking directly against anthropic by name on token efficiency not capability
is the tell that 2026 is model war is now a cost war for two years the story was who's the smartest
but now it's who's the cheapest per agentic task and that's a different competition with different
winners. Put it next to what we just covered, cursor getting similar quality at a fraction of
the cost by restructuring context instead of adding agents. Add Opus 5 explicitly positioned
as the most of Fable 5 at half the price. And you can see the whole industry converging on the same
axis at once. I think every lab ships an efficiency benchmark against a named competitor within the
next two model releases because Altman just proved it's a headline, not a footnote. What do you
think, Henry? Is efficiency bragging just marketing theater until independent benchmarks confirm it?
HiM
Yeah.
AndyML
Or is this a new axis? HiM: I mean yeah I mean I think for Anthropic it is right I mean um you know I'm not a fan of
HiM
Anthropic but I think for Anthropic it definitely is marketing um I don't think they care about
decreasing token efficient making tokens cheaper because because that's
the reason Anthropic is is so far in terms of revenue is a few things first of all their
revenue isn't as far as it is they count gross revenue so the revenue for anthropic that is
reported is gross revenue not their actual so like net revenue open air accounts net um but
anthropic the reason they make a ton of revenue is because their tokenizer is always more expensive
so their models are very expensive so i my pushback would be that for anthropic is definitely
AndyML
Yeah.
HiM
is marketing um open air has always cared about costs open air models are always like cost efficient
um so they care about cost um opus 5 is only 20 cheaper than than than fable but fable is like
mad expensive compared to like everybody else so 20 cheaper it's not it's not that much
AndyML
Oh, yeah.
Yeah.
HiM
like it's mad expensive right can you can you see those like i think you only have um you only
have anthropic models but if you were to compare with like open ai and like keemi it's ridiculous
loss it's it's so expensive uh so yeah so for me i would think that yeah it's and for open ai it's
definitely kind of like part of their they believe it and and they've worked very hard to get it
there for anthropic this they just don't care about it they just want to make as much money
as they can from enterprises um no matter how expensive it is enterprises are going to pay
even though i suspect that the reason they're doing the marketing now is because um kimi is
Kimi is out
you know
out and swinging for the fences
and it's like
10% of Anthropics
like are feeble right so yeah
so now I think that they're gonna care
a little bit about it but I think right now
it's still marketing
AndyML
yeah i mean 5.6 soul is 30 per million tokens out compared to claude opus 5's 25 or
fable fives 50 yeah it's just not even comparable
HiM
yeah anyways um so yeah so um
AndyML
interesting well i i think we'll end up seeing you know what people end up using
HiM
i mean i'm not a cool fat i'm not a cool fact for this as well right i mean it's like they're just AndyML: um as their strategy right
AndyML
All right.
HiM
so far behind like guys i just do just take them a while to catch up right i think this is probably
you why they also did opus 5 because it's pretty much feeble but just cheaper um so it's a topic
that people care about um it was in the news today that grog 4.5 just passed um all the other
frontier models and asked to become the top used model and an open router so it's now it's not used
the most tokens and open router um and the reason is is the point you're making and it's just way
cheaper. Correct. AndyML: if it does the job and it's cheaper.
AndyML
So someone mentioned, was that Trev?
Someone mentioned the mixture of experts
is going to end up being like an orchestrator layer, right?
You're going to use Grok 4.5 for this and Sol for that.
And maybe you'll use Fable for a plan.
Maybe you won't.
HiM
Yeah. AndyML: You'll probably do an A-B test a couple times
AndyML
and then settle on something else.
HiM
yeah and um and um and cursor launched the router this week right so AndyML: Yeah, interesting.
AndyML
Yeah. HiM: cursor launched the router this week pretty much back the same idea right so people are people are
Yeah. HiM: starting to care about so for me i agree that cost and efficiency is important it's important
HiM
to the enterprises um but whether the labs like anthropic are taking it serious or whether they
have the ability to catch up is what i'm arguing against but i think for enterprises cost and
efficiency is important to them which is why they're going to use open source which is why
they're going to start using chinese models which is why they're using grok um so it would take a
for the guys that come through to catch up.
AndyML
yeah fair enough reasonable take be an interesting one to watch
um this episode is also brought to you by heritage telecom uh trusted phone systems
from trusted people heritage designs and supports the whole communication stack for your business
dependable phones, failover, reporting, and practical AI that turns calls into action.
One accountable provider who actually answers. Start with the problem, not a product list at
heritagetell.com. Well, a few things to watch heading into next week. Keep an eye on whether
OpenAI publishes more details on the Hugging Face incident. There was an actual zero day,
HiM
I'm going to go ahead and get it.
AndyML
and so I'd be curious to see how it was patched.
We caught it isn't exactly the same as it can't happen again.
Second, watch whether real teams actually migrate off of Slack
or GitHub onto Buzz versus just kicking the tires on a developer preview.
And third, now that Opus 5 is the default on Claude Max,
watch whether Fable 5 usage holds up in the
I need the absolute best option sort of a way
or whether most people just stop needing it.
um we sort of feel like you might not need it so everything from today sources the transcript
HiM
Yeah.
AndyML
henry's research sheet some show notes um it's all available at weeklyclaw.ai
full episodes and highlights on youtube clips and takes will show up on x and we're back next
friday july 31st 4 p.m eastern so follow the excitement and we'll see you then
HiM
yeah we should um we should we'll probably do a poll next week or put a poll on the website
so you guys can let us know um what's a good time is this still a good time to do this show
um or if you should do it at another time we talked about in the beginning of the show
AndyML
yeah right on we did and i think you've got in here where is the feedback link
HiM
Yeah, the feedback link should be on the bottom.
AndyML
oh here it is so at the bottom of the page give audience feedback i would encourage all of you
guys to visit weeklyclaw.ai and give us some feedback thanks what skills appreciate you
but yeah give us some feedback we'd we'd really like to we'd like to know um you know what's HiM: Awesome.
good what's not like to constantly improve like our agents so anyway thank you all
HiM
Yeah. HiM: let us let us know um let us know whether you like the agenda again we're kind of like uh moving
this around um but let us know if you if you like the agenda and if we should like add something is
AndyML
yeah
HiM
there something you would want to know um or you'd want us to talk about um but if this is good
we'll just keep it going all right
AndyML
yeah sounds great awesome well thanks so much everyone have a great rest of your day