Weekly Claw · episode 30
18 Sept 2026
Narrow Models Win, Typed Decisions, Union Alpha, Signal From Outside
What this episode covers
- Narrow models lead the week: typed probability outputs and cheap decision endpoints beat another general chat model for operator work.
- Union Alpha appears as a free 256K-context stealth model on OpenRouter and becomes the access story of the episode.
- Qwen and PrismML merge into a practical omni/flash lane while Apple’s Siri AI beta and Google Home MCP open the assistant into the home.
- Signal From Outside covers boring good news already working in the field, from medical and sensory aids to storm forecasting.
- The closing debate asks who pays for an independent safety umpire when commitments are not controls.
Published record
1419 published segments
Speaker not identified
Here are three places where it's already
happening. Each one has the same pieces
we build with every week, right? Inputs,
model, action, and a human who makes the
call.
Let's fix now. Um, I think it was a
it was a good week. It was like a it was
a humanpace week, right? It wasn't
crazy. Um, I like it. I actually like
this week a lot. Um
especially because of the kind like the
company that came out of stealth with
the product that was like really good.
Um
so I think it was a really good week. Um
I think I was productive. Um I mean last
week I was saying I wasn't wasn't as
wasn't as productive as I would have
wanted to be. I think this week I was as
as as um it was this week is definitely
much better. Um so I think an AI land is
a really good week. It wasn't too fast.
Um, not a ton happened. Um,
[clears throat] or maybe a ton happened,
but it's just that we've gotten so
desensitized to to to all the um all the
things happening.
>> Yeah. Right. Launched
>> and Yeah. It wasn't a big week. Like, it
might just be changing the history of
mankind, but it also might not be a big
deal. We'll we'll see.
>> Yeah.
Yeah. Oh, God.
>> Love it. Well, good.
Well, welcome back to the Weekly Claw
episode 30. Here we are. Um, Agentic AI
continues to um amaze and astound
even in a slow week.
You know, there's still some pretty
exciting stuff. Um, some applied AA
product lab
products. Uh,
we've got
I Yeah, I mean there's news and there's
is I I'm going to try something this
week with the signal from the outside.
Um, instead of just reviewing sort of
one video,
um, I kind of came up with a theme and I
put some pieces together from a few. So,
I hope that's enjoyable. We've got kind
of a shorter show today, but um, I think
you'll enjoy it. So, we'll try not to
waste any time.
Uh we've got type safe obviously. Um Jev
is bringing some big changes to uh I
think pretty much everything. I think
people are going to try to squirt it all
over the inside of their carry out bag
um and add more dev. Um
Apple bringing Siri
license from Google. That's exciting.
Uh Mozilla is going to let you pick a
model. Yeah. So, I mean I we don't need
to overview all of it, but um it should
be a a decent episode. You have anything
you want to say before we get started?
>> All right. Well, before before we do,
this episode is brought to you by Herold
Labs, an applied AI product lab where
humans and agents build together. Entity
is mission control for agent teams.
Hacker houses worldwide. Build with
humans and ship with agents
at Herald House.
Labstheold.co.
Right
>> there you are.
>> There we go.
>> All right. All right. Um, so yeah. So
guys, I mean, it's not, like I said, it
wasn't a I mean, a lot happened. Um, but
we try to just like select the most
impactful. I mean, there's always a lot
happening and we're pretty much in a
singularity. um the piece of the piece
is always like incredible. So there's
always a lot a lot happening. Um but I
selected the most like impactful news.
Um I mean I don't think I have it on the
slide but I mean I mention it I mean
towards the end of last week um there
was this scare about you know the
anthropic exanthropic employee that when
AI is going to kill all of us. I think
we talked a little bit about it last
week. And then in the weekend on Sunday,
Dario writes uh again one of his long um
articles to say, "Hey, let's let's pace
the frontier um and let's um
um let's add human let's add third party
evaluators to to each [snorts] like
company." Um so that happened that's on
the slide because I mean um you know so
that happened and then obviously that
led to kind of like a lot of things. I
mean this is sort of like what will I
call it? It's it's not really kind of
like an event. A lot happened. Elon's
like yes. Sam was like yes let's do it.
Um De was like yes I agree. Um Zach was
like nope I don't agree. Um the
president was like nope. Geeks I'm not
slowing you guys down. China is like at
our heels viewed on um and then there's
like all this like um
um back and forth different camps we
want you to regulate we don't want to
you regulate um I guess we're going to
talk a little bit about it and and
signals from outside but yeah so a lot
happened um there are a couple of yeah
conferences obviously AI themed there's
all there's all in summit there's the
Salesforce
um Dreamforce and they're all AI themed
and there's so
interviews and so much content. Um,
obviously Andy is going to kind like
talk a little bit about that. So like
again I didn't put those things on the
slide but there's a lot that happened in
that. So like theme
um so the most impactful what I think is
the most impactful news from like this
week
um is obviously Jev um from type safe.
Um so Typescafe is again an exopai
um executive
um
he was the co-founder of charg as they
like to say um or he likes to say um and
he is no sort of like the kind like
founder or co-founder of hlf which is
like this um
this process for kind like training
models or doing post training
um And so he's come up with you know and
this isn't new just that he's launched
and so we're hearing about it today uh
well not this week but you know he's had
this company called type safe I actually
watched a video that he did like a few
months ago at AI engineer summit and
he's basically talking about this idea
right so every basically what they've
launched he talked about it months ago
I'm sure he's been talking about it for
a while but until people see the product
they don't typically kind like grock all
this concept so he's been talking about
it for a
um which is what he calls an RHCD
something like that um reinforcement
learning
um decision something decisions right so
he's basically launched a model that
isn't necessarally a large language
model or even if it is a large language
model it doesn't generate text um it
does u it generates decisions right and
so I watched his video and and is going
to play the launch video for us and
maybe you should play that video before
I like and te te tease out like some of
his ideas a little bit. [clears throat]
>> Sure, let's do that.
>> I'm Diego Almeida, founder of Types Safe
AI and at OpenAI, I co-created Chat GBT
and RHF, the post training algorithm
behind most frontier AI. Our team
trained the first models to be
superhuman at instruction following,
[music] what we now call chat. And we
asked ourselves, are models that are
superhuman at chat AGI? The answer was
obviously not, but the trillion dollar
question is why not? RHF has led to LMS
that are optimized for human preferences
and include issues such as mode
dropping, overconfidence, and an overall
lack of reliability. These flaws mean
that LMs require humans in the loop, and
almost no true automation can be done.
At Typesafe, we've spent two years
building in stealth and we're finally
ready to share our new type of
foundation model that's optimized for
automation. System one models with a new
architecture, new sampler, and new
training algorithm reinforces learning
for calibrated decisions. The
improvements are clear if you see them
side by side. Ask a system one model a
ton of structured questions just like
you would an LM. [music] Get the answers
back near instantly. Meanwhile, LMS take
hundreds of times longer to finish
responding. Look at how the LM generates
sequentially, which is great for a
natural conversation, but totally
useless for computers. LLMs extract
intelligence from the tiny straw of auto
regression. And similar to the jump that
transformers made over RNN's, we are
replacing sequential computation with
parallel because that [music] is
obviously the future. Our system one
models output decisions with
probabilities and confidence instead of
words. And they can't hallucinate.
They're a lot more like code. Reliable,
[music] fast, self-consistent, and type
safe. This opens up a whole new world of
possibilities and applications. Today,
we're releasing [music] Jet, the first
public system 1 model. It's 100 times
faster, so realtime AI is finally
possible. 100 times cheaper with input
tokens priced at $42 per billion tokens.
and Alpha tokens are free because
they're finally too cheap to
meter. Jev is super smart and its
intelligence per dollar is literally off
the charts. This is just the beginning.
We're excited to see what you build with
Jeb. As we say it at Safe, we're
building prod, not God.
>> Yeah, super exciting, right? Um,
so I'm not sure how many of us have been
have started using this model um from
the psychic community, the guys watching
us live. Um, let us know in the chat if
you've if you've really started using
the the model. Um, but it is um it is
incredible, right? Um, Andy has started
using it. Maybe you talk talk to us a
little bit about it when I'm done
recapping a little bit. I have I haven't
started. And I mean I I've used I got
access I got access through Versel AI
gateway. Um but I [clears throat] guess
like the like the most impactful from
his video that he did on an AI engineer.
He's talking about how the reason he
wanted to do this company
um [clears throat] uh was that he always
like had this dream or what we were so
this this dream of like software being
like in like super intelligent and super
smart. Uh but with LLM software isn't
necessarily intelligent or smart. We're
building software faster but the
software itself isn't as intelligent
right um there are very few people who
are embedding LLMs into their software
even the ones who are I mean it's it's
so expensive that it's at scale is still
not really reliable to be embedded like
in almost everywhere so he's so like
vision is to have an AI model that can
be reliably embedded across all stacks
of software and it just makes the
software smarter and you can think about
and it's almost
um which is what they've basically done.
Um what else is valuable to mention?
I think that's pretty much it. Um so
it's a great model. Um there's been lots
of videos. I've been watching lots of
demos. People have been building like
compaction which is one of the things
and he tried um browser use is like off
the charts.
Um the the every guy every team did a
very long article um of very different
use cases. Um, and it's just incredible.
It's just an incredible model. Um, and
so just to recap for it, I am very
excited about this. Um, and I'm hoping
that, um, for instance, the last the
last phrase, um, or the last like
tagline they have, we're building
Prague, not God. Um, is again coming
from all the drama that happened in the
weekend and last week. Um, the frontier
models are sort of like drama queens at
this moment. Hey, we're going to kill
everybody um because we're building this
thing and we don't know how to stop. Um
and it's so like right? And
I'm glad that a company like this showed
up the next week because this isn't a
high-risk model at all and it pretty
much takes over almost 50% of all the
use cases that anybody would want to do.
And so I expect this to take like market
share from the LLMs from the kind like
frontiers. I mean it's not a frontier.
It's not a it's not a proper LLM. You're
not going to chat with it. you're not
going to generate review with it. Um,
but you can start to do very interesting
classifying work, decisioning
[clears throat]
um and and my kind like companies have
already sent to the engineering team
like hey let's start deploying this
thing um immediately. So I hope that
this takes a ton of market share from
from select the existing companies. Um
because then it deflates the ego, right?
Cuz there's a lot of ego right now
around like we're beating God and we're
beating this thing is going to kill
everyone. But to be honest, man, if
there's no economic economic value in
beating AGI, no one is going to beat it,
right? So if if we've got a model like
Jeff that is almost free and can do half
of what um the models that all the guys
beating AGI trying to build then um I
think the industry will be a lot safer.
So that's that's why I'm super excited
about it. So I hope that pans up. But
Andy I'll let you respond and I'll run
through the rest.
>> Yeah. No, I mean you put your finger on
it, right? A lot of industrial use for
language model is essentially theory
rigging the language models to produce
structured output to get at the very
things that Jev produces natively. Um
the number one spend on my portfolio is
running call transcripts through um call
categorization routines and then
depending on the category of the call um
deterministic scoring. I mean, it's it's
it's inferistic deterministic scoring.
And you know, Jev can probably do it.
We're we're we're almost entirely sure
we can take the scorecard for the call
with the transcript and then just
literally use the probabilities if we if
we were if we word the questionnaire
questions, right? Like this service
advisor
said the name of the company in their
briefing. Like probability 98% because
they said it. probability 2% because
they didn't and you know literally use
that as our scoring. Um it's going to
save us at production across
uh 350
you know sort of paying customers on a
particular platform
maybe $2500 a month just right off the
bat right so we're testing it as quickly
as we can um I think there's a great
opportunity I compaction is an
interesting it's going to need to be
part of compaction it it can't do all of
compaction And um the implementation
that I tried was sort of halfbaked and
it it tried to replace the compaction
layer without any summaries and um I
think if it were used to protect the
work of a compaction layer uh it would
be very effective but I don't think it
can just replace it.
Um and then the other piece of course
we're all playing with is trying to use
it to route requests to the right model
based on the request. um you know what's
the probability that this um request can
be handled by this model with these
capabilities
um you know without making mistakes
right 30% 40% 80% well now we we know
which model to use for this request and
then if it's a simple one it'll you know
we'll get a a cheaper model so anyway
model routing and um categorization just
out of the box very excited to see
browser use Um, you know, excited to see
OpenClaw, the G OpenClaw community, the
Hermes community, they're all going to
build, you know, browser use plugins
that leverage it. It's going to be a
major component in all of these
harnesses before long.
>> Yeah, it's incredible. Um, obviously, I
expect, yeah, like I said, it's not
going to be a replacement for for like
LLMs. Um but it is um I I expect and I
wrote a tweet about this. I expect that
the other companies will eventually
launch something. Obviously China will
launch something super soon. Um this is
a totally new use case um or totally new
like form factor. So I expect the the
frontier model companies to also launch
their own version just so that it can
keep customers from turnurning off their
platforms. Um, but it's so cheap, right,
that even if they launched equivalent
like products, they would have to like
match the same price and this price is
just like just eats their launch. Um, so
it's exciting. So again, um, this is the
biggest news um, in my [clears throat]
books for the week
or in my book for the week. Um, another
interesting thing that happened, I mean,
it started earlier in the week, um, was
that, uh, a stealth model showed up in a
few of the model market places called
Union Alpha. Um, Union Alpha was pretty
good. Um, was a 26 262K kind of context
um, model. It was is multimodal. Uh,
it's pretty good, right? Um and I think
in a in a in 2 days or so um the like
community or kind like the world really
spent something coming close to like um
um was it like now [clears throat] it's
like 100 billion in the first day and
then I think it's closer to like 600 700
billion something um tokens. I think it
was close to 1 trillion tokens really in
the two or three days that it was on.
It's off. It's It's off now. It went off
like yesterday. I was trying to use it.
I used it for a couple of days. It was
pretty good. Um it was pretty good. Um
from the demos, um Andy, I'm not sure if
you used it at all, but from the demos I
was seeing is you didn't use it. Okay.
From the demos was um it was really
good. It was pretty good. I I put it I
put it to work. I mean, obviously
because of rate limits, I couldn't
really benchmark it properly. Um but it
was really fast. Um I could, you know,
maybe stream like north of like 100
tokens per second or something even like
double that. It was pretty good.
>> Real fast.
>> Um, yeah, it was really fast and it
worked. I mean, obviously see for free
models, um, you get really limited. It
wasn't very reliable, but it was really
good. Um, and so [clears throat] the
word on the street, um, and obviously
when the ST models show up, people sort
of like start to kind like um, you know,
throw out their their theories around
which company owns the model. Um, so
they think this model is is an open air
model. They think this is probably GPT6.
Um yeah.
>> Oh.
>> Um so I mean obviously people know how
they kind of like people know how to
like find the things by like comparing
the quality of the work he does like
other. So yeah people were also doing 3D
um you know 3D games and stuff with it
and and it was lacking a lot of soul and
we know that OpenAI most likely will
launch um GPT6 soul um at their death
day next week or something like that. Um
so yeah so um um fingers crossed we'll
find out um what model it is hopefully
next week or on the week week after that
um and then the model launching um model
launching kind like news um two other
models dropped this week um Prism ML
launched there again the Bonsai guys um
um they typically would fine-tune the
queen model the queen models um and so
each time there's a new queen model
you'd expect that they would launch
their and it's like fine tune of that.
Um and [clears throat] so the the
finetune for the 3.827B A27B model came
out um which is what they call the
Bonsai
um you know 22 27B right? Yeah, 227B.
>> Um, which again you can kind like see on
there.
>> Um, it is it's like a two bit model. Um,
right. It's 262 um K of context. Um, and
then they released um it's Apache
license. There's a GGUF and MLX and then
a CUDA. Um, and and so I mean the the
exciting thing about this, why am I
telling you about this, right? And I
mean, obviously, if you guys um if you
watch the show, you would remember that
um one of the models that I've been the
model that I've been the most excited
about this year has been the Quinn
3.827B. Um because it's a really small
model, it can fit into most laptops. Um
and then it comes with like pretty much
like Opus 4.8
kind of like quality, right? And
>> it's real good.
>> Yeah. I mean, Andy uses it, you know,
sometimes. Andy's is is off to Bonsai
now. So maybe he's also uh a good person
to talk to us about what the experience
has been so far. Um I haven't set it up.
I'm I'm benchmarking it at the moment.
So I don't have it live. But obviously
what is exciting about it is um this is
a quantized version of that model. So
it's a two-bit model. Um the 27B is
obviously something in the universe of
27 billion. Um
um but this is five, right? It's it's
it's really small. It's like nine times
smaller than a typical model, right? Um,
and there's almost no loss in in quality
or performance. It's like a 2% loss in
setting benchmarks, but it's it's it's
almost exactly as as um the model that's
nine times its size. Go ahead, Andy.
>> No, I I mean, I just I just have to
point out, right, it's it's 98% as good
as Quen 3.827B, 28 27B and it's nine
times faster and it uses like um I
benchmarked it on just just you know
real high-speed low drag tests um an M2
Pro uh Mac uh Mac Mini and an M4 Max
um MacBook Pro and it absolutely rips.
It's very fast. Um I did benchmark it.
You'll be interested to know Henry
against Ornith 1.5 which you know I'm
very excited about 35B model came out
about maybe 3 weeks ago and um it ornith
outperformed it in coding and tool call
tool calls but only barely. Um and the
performance is excellent.
Um, Bonsai was doing 70 tokens per
second on a MacBook Pro uh, M4 Max and
maybe like mid 20s, 15 to 20, 25 tokens
per second on an M2 Pro. So, it'll run
on modest commodity hardware at usable
speeds. It's smart. It writes good code,
and it's small. It doesn't even take up
a ton of your disc space. It's not like
you're using tricks to load part of it
from, you know, stream experts from SSD
or anything. It's just all right there.
And it loads on a 16 gig GPU. So that's
what the 4060 470
and up. U very fast, runs very fast on
those. So any modest GPU and you've got
a very powerful um model to run your
agent
>> um run a model. Um the models are
getting good like this is this is um
this is um 54 just 6 gigabyte so 16 gig
should work um so
>> um pretty much all right cool the other
thing that happened was um was the queen
guys launched a 3.8 Omni flash. And I
mean again, why this is um why I guess
this is interesting is we're starting to
have um the model companies or the
Chinese model companies now launch flash
models. Um they all used to have two
model classes. They would launch a pro
model and then a flash model. And then
the pro model would be like the the
heavy hitter and then the flash would
like be that's like lightweight. Um but
we're starting to see that if that
there's a merge coming where they just
launch one model um that is pretty much
like the best that they can, right? Um
we're seeing this with Deep Seek. This
Deep Seek the recent launch of like V um
4.1
um you know it's called is it called
Flash? I'm not sure that one has Flash
in it. Did it have Flash in its name? I
don't remember if it did. Yeah. Okay. Um
so [clears throat] yeah so so we're kind
like having this trend where we most
likely soon would no longer see pro
models from the Chinese. We just see one
model which is again size of a typical
flash model but with a performance of a
pro model um which is what people want
right people want like models they can
run cheaply on their hardware that is
frontier performance. Um so um so 3.8 um
the last 3.8 8 model um that probably
has this like benchmark so it's closer
to the 3.8 Max um if you remember that
was launched at the same time with
3.827B.
So with the 3.0 or mini flash um we're
pretty much getting again super like
cool performance at like a slightly
lower um size. Okay. So um I mean
obviously there are a few other things
that happened and like the open source
community um but these two were things I
thought were super important as like
mentioned to you guys. Um um again
obviously I don't have this video but I
remember seeing a video I tried to find
it ahead of the show but I couldn't find
it. um was that I mean [clears throat] I
have a few friends that are on the Apple
better program or the MacBook better
program or Mac [clears throat] OS better
program and so they've been getting
updates to um the new OS and they've
been pretty excited. Um this week um I
think it kind like came out as well that
you know Siri is in better now. Siri AI
um anybody
everybody's kind like sees that Apple
the bed a little bit with Siri,
right? Siri was like the perfect form
factor for like personal AI, right? It's
like um
>> had everything [clears throat] it
needed.
>> Correct. They had like billions of
users. It's it's a mobile. It's on the
PC. It already has the voice form
factor. Um but but yeah, but but the the
Mac the the Apple just didn't get it
right. Right. Um [snorts] um but it
looks like they might be um they might
be coming back. Um who knows? Um so
people who've been using Siri AI that so
far I've been seeing they're very
excited about it. It works really well.
I'm not sure if anybody in kind of the
live community has it or they're using
it. Um if you are, please let us know.
Um but yeah, but but you know, I haven't
used it. Um I don't I don't use the
better I don't I'm not subscribed to
Apple's better program so I haven't
tested it yet but I thought it was worth
mentioning. Um if you also remember
[clears throat] Apple partnered with
Gemini or with Google last year and and
you [clears throat] know their models um
kind the the products will be powered by
Gemini models um hosted on Apple
infrastructure. So I'm expecting that
that that that Siri AI will be powered
by by Gemini collect model. Um, and
Geminina has been pretty um I mean
obviously that that leads us to like the
um final like news on my docker. Gemini
has been pretty um pretty good with with
um voice models. Um there were a few
voice models that came out this week. I
don't have it on on the slides cuz they
weren't as super important. Um but there
is there are a few voice models um that
they put out and they're pretty good and
they're not as good as GPT live one from
OpenAI but but they're they're pretty
good.
Um, and then finally, um, I I I I
thought it was valuable to kind like
mention the MCP, um, you [clears throat]
know, the MCP access to their hardware.
Um, um, because I I have a friend or I
have a friend actually, yeah, that that
wanted to,
>> [clears throat]
>> um, put agents in in his assistant like
the um,
[clears throat] what's the Amazon one
called again? The Amazon hardware.
>> Alexa.
>> Alexa, right? So he wanted to put like
agents his agent on Alexa or on the on
the on the Google kind like hardware and
couldn't do it. So, um, and so he had
Astraat teach him how to build his own,
um, speaker, like smart speaker, so he
could build a smart speaker from
scratch, uh, and so he could install,
uh, his own like a gen into it and so he
could just talk to it, right? And and
and but but this is cool because I mean,
um, they're starting to get there.
They're starting to open up these
devices. Um, so with MCP access now, it
means you can be pretty much have an
agent um, use tools and you know, maybe
you can send them stuff to play on
there. You can have them control like
things. So, we're getting there. I
probably get to a place where I mean, I
don't think we'll ever get to a place
where they would let you load your own
agent. They have their own agent and let
you access it through that. But, I think
we're we're kind of like making
progress. I think that's pretty much it
for me this week. Um, there are a few
other like interesting things that
happened. um Anthropic um followed um
cursor uh launched projects. So now you
can you have projects um and you can
talk to one agent and then he manages a
bunch of sub agents um and um yeah I
think that's it for me and take it. Oh I
know I know something super exciting to
kind like mention um I mean two okay two
I mean OpenAI have actually been on the
news this week um quite quite more than
they should. Um, OpenAI obviously
released like a lot more like
documentation on a lot more kind like
hacks um and and and stuff and the
framework for reporting new hacks. Um,
Anthropic two news from Anthropic today,
right? Um, one is what I was just
telling Andy just before the show
started. Finally, Anthropic is going to
do agents. MD. So, you no longer have to
use clot. MD. Um, I think starting from
a new version that launched today. Um, I
mean people have been like um talking
about it with them. was like, "Why don't
you just use the same standard as
everybody else? Why do you think you're
special? That you shouldn't do that." Um
so I think today they they finally um ti
just announced it a few minutes ago that
now going [clears throat] forward like
agents empty, you can use agents empty.
Um and then a much more scary news which
again I don't think I mentioned to Andy
um but I I just saw that um it was just
announced that Anthropic started quietly
started a lab. So they have a a biolab
and they're trying to make um Yeah. So
that's the thing I saw. Um
um so yeah. So
>> making a bolab.
>> Yeah. So that's been um so that's super
super scary, right? It's like super
scary, right? So everybody kind
like I know in the community and they're
like they're like no man. Like um so
yeah. So
>> yeah, man.
>> Warning us that AI is going to kill us
all.
>> Yeah. Yeah. Yeah. Yeah. Yeah. So that's
um I mean those things tend to be you
know Elon always jokes about it right
that people tend to kind of like be like
opposite of what they so yeah so I think
that this is actually I mean obviously
I'm the I'm the I'm the don't regulate
just chill because there's enough
regulation right now but I think for
this one I think it's actually they
should probably stop them right because
I mean um these guys are uh I said it in
the group one of the groups I'm in right
these guys have the highest speed doom
um of any company or any org um and
they're definitely not the right people
that you want to like be getting close I
mean sure make software right make
software make AI that's fine that's far
from like the real world but like when
they start getting close to like giving
like models access to like bio equipment
then we then we know that like like
yeah someone cuz man all the risk we're
talking about like AI AI AI only those
labs have enough computing
damage right
>> exactly right the other like hugging
face and me and you can't do it. We
don't have like enough compute to like
send a,000 agents to hack someone,
right? We don't have 10,000. We don't
have enough comput like send 10,000
agents to do any work, right? Those
things cost like millions. Like um so
yeah, but but the labs do. Um and so
yeah, so when you have a lab um start
playing with bioweapons or start playing
with trying to make drugs, no no guys
are going to now I can see
happening and heating [laughter] the
>> Yeah. Yeah. Hopefully they're not just
trying to make their claims true.
>> Yeah. So that's it for me guys. Cheers.
>> Yeah. [snorts] Hey Henry, thanks for
that. It's uh good overview. Interesting
conversation on a somewhat boring week,
but it's interesting that uh even when
it's boring, it's not that bad. Um look,
this is my um my signal from the
outside. I am
I don't know if I need to apologize for
it up front. This is not what we you
typically do. Usually we'll focus on
videos that are um really relevant to
Agentic AI uh builders, right? That's
our community. Um that's Henry and I and
that's um hopefully very many of you.
But the news this week was discouraging
enough and I I even found myself in a
couple conversations where I didn't have
some answers that I wish I had had um
against sort of people watching the
news. So, I thought I would put together
a little bit of an aggregate of a
handful of stories. Um, so this is the
signal from the outside. Uh, like I
said, it's been a heavy week. If you've
had the news on, you've heard Swarm or
Botnet or Take Over the Internet more
times than you'd like. And they're real
questions and serious people are working
on them. Um, I don't want to wave that
away, but this week, two very different
people pushed back on that fear and they
ended up in the same place.
uh one is the most important supplier in
all of AI and the other is a finance
writer with nothing to sell whatsoever.
So I want to start with them and then
spend the rest of the segment on what AI
is already doing for the people outside
this room and I'm hopeful that it will
provide some perspective um for our
conversations with those people. Right
on Monday, Jensen Hang sat down at the
all-in summit. He was asked about the
week's essays and the push from the
frontier labs to slow down. He started
by saying safety is paramount and has to
be taken seriously, but also that safety
versus leadership is a false choice and
that you can move fast and do it safely.
And then he went after the forecasts. Um
his case was track record. He used
radiology as the example. Of course we
all know years ago uh the prediction was
that AI would take over radiology and
there would be no radiologists left. Um
what he said literally was uh that has
proven to be exactly the opposite. Now
we need more radiologists than ever. And
in the same breath AI now does a huge
share of the scan reading which he calls
great. So the work changed the people
are still needed and the patients get
read faster.
That's kind of a net positive. It's not
like we have a bunch of unemployed
radiologists.
So he ran through the list, right? Most
code written by a AI within months, half
of entry-level jobs gone, models too
dangerous to release. His view is that
those predictions haven't held up, and
that somebody should be keeping score.
You can take or leave Jensen's position
on policy, and plenty of people in this
server will disagree with him, but the
radiology point is worth keeping because
it's the pattern for the rest of the
segment, and it seems to be the pattern
for the track record of AI. The second
voice is Morgan Hel, who writes about
money and behavior. Last Friday, he put
out an episode called AI optimism and
the agony of waiting, and his point is
about us, not the technology. what
you're dreading. When you're dreading
something, he says, your mind is a very
proficient storyteller. It fills in
every question mark with the worst
ending. And what you can't picture at
all is the ordinary useful stuff that
happens all the time. He offers one
forecast and says it's the only one
he'll make. I think it's a little
conservative. He says, "10 years from
now, the likely story of AI is pretty
mundane. a lot of jobs get 20 or 30%
more productive and the gains go into
growth, wages, and cheaper products and
life will go on. He brings up the late
'9s.
Uh, if you remember the late '9s, people
were sure the internet would wipe out
retail jobs. He points out that there
are more retail workers today than there
were then. And two of the biggest
business successes of the last 25 years
are Costco and Walmart. Of course, they
both have online presence, but those
facilities are open. You can go to a
Walmart 247 in most cities.
That's Jensen's radiologist again. The
work gets better and the people are
still here. Carrie Herrell uh Harrell, I
I assume um Casey Harrell is an
environmental advocate in California
with ALS. By the time he joined a study
at UC Davis, his speech was very hard to
understand. Surgeons implanted small
electrode arrays in the speech areas of
his brain. When Casey tries to talk,
those neurons still fire. AI models
decode the patterns into words, and the
software speaks them in a voice rebuilt
from recordings of Casey before he got
sick. It sounds like him, and the first
time it worked, he was talking to his
family within minutes. In June, the team
published two years of results in Nature
Medicine. more than 3,800 hours of use,
close to two million words at about 56
words a minute, and 99% word accuracy in
testing. It runs at home, operated by
his own care team with no researchers in
the room. He's back at work full-time.
In a UC Davis video last week, the
researcher calls him the ultimate power
user. That's the bar for an agent. It's
not a demo that works while you watch
it. It's thousands of hours unattended
for someone who depends on it. The
second story is the most us of the
three. Um, one disclosure, it does come
from a video that OpenAI produced and so
it's partly marketing.
Uh, Ryan Honory built a heat detector
for his fifth grade science fair.
He kept it going for years and it became
sensory AI, a network of sensors in the
hills above Lagona Beach that try to
catch wildfires before they ignite. The
sensors detect a language model turns
messy readings into a plane alert. Ryan
built a walkie-talkie interface so a
firefighter can hold down the button and
say, "Why do you think this is a fire?
What are the GPS coordinates?" In the
video, he also has it schedule a sensor
test every morning at sunrise.
Sensors, a model that translates, a
scheduled task, a voice interface, a
human who decides whether to roll a
truck. That's the agent stack pointed at
the ridge line instead of an inbox. As
one of the fire officials in the video
says, "If a fire gets large, it's
unstoppable. Catching it at ignition is
the whole game, and minutes matter."
The third is the biggest in scale. The
source of course is Google DeepMind and
they're describing their own model. So
take that as you will. Last October,
Hurricane Melissa hit Jamaica as a
category 5. On a podcast last week, pet
uh Deep Minds Peter Betiglia said that
almost a week before landfall, their
model became confident that the storm
would reach category 5 and it was barely
a tropical depression. The model didn't
issue the warning. The hurricane center
did. Their forecasters weigh a lot of
inputs. Traditional physics models,
newer models, their own observations. By
particularly his account, the center
told him uh the center told them the
model's confidence raised their own
confidence, but it didn't change all of
their math. They made the category 5
call 3 days out, which he says was the
weakest storm they'd ever forecast to
reach that level. Melissa was truly
devastating, but three days of warning
means evacuations and preparation.
Same pattern as the radiologists. The
model was just one strong input, and the
humans with the authority made the
better call earlier because of it.
Here's what the good news looks like
this week. A man with ALS talks to his
family in his own voice and goes back to
work. A teenager sensor network watches
a hillside and answers firefighters
questions over the radio. forecasters
get three days of warning on a monster
storm. None of these is a takeover.
They're mundane in the best sense, like
Jensen's radiologists. Each one is a
model doing a narrow job well inside a
system built around people. And the
human stays in charge of the decisions
that matter. That's the same thing most
of us are building. Our agent that files
tickets or triages our inbox or watches
our logs at 3:00 a.m. isn't going to
make the evening news. But if households
right, the real story of the decade is
millions of builds like ours, each
making someone's work quote 20 or 30%
better. The fear is loud this week, so
keep building the boring good stuff. And
that's the signal from outside.
>> Thanks.
>> I mean, it might be worth it.
>> It's worth saying.
>> No, I mean, I think I think I think it's
good to like getting to the habit. I
mean um this even if um I mean we don't
have a lot of um I mean the community is
growing. I think cumulatively if you add
all the channels maybe we get like 200
views a week or something like that or
something someone would hear it. It is
valuable. No no of 200. I mean Jim was
telling us that the Chinese platform
gets close to 100. So so yeah you could
say maybe 200 300 right when you can
like piece up Twitter and all the other
places. Yeah.
>> Um so it's valuable right? We need to
and obviously um we just like mentioned
it and we'll probably do a proper launch
sometime um but Bandandy and I were
talking about this um yesterday and it
was an idea I had for for a few weeks
now not going into months. Yeah. Start
to talk about a bit of um you know
optimism, right? There's just a lot of
[clears throat] doom out there and
nobody's telling the story that Andy
just stored and we decided to put
together a thing and so we've got a
website and we share with you guys super
soon. We put the domain yesterday and
and we've put out um um we've put up the
uh the select uh gender and and it's
just been hey let's tell the optimistic
stories right let's let's get people
excited about like cuz AI has been a net
massive net positive in my life it has
been for Andy's I mean this whole
podcast is basically run by agents right
like if we didn't have the agents and we
couldn't be doing this right like it's
just a lot of work we've got many
businesses to be running nobody has time
but so the normal people would could be
empowered, right? So, um, so instead of
listening to all this from
people who are trying to capture power
and then making you think, hey, this
isn't, you know, no, it's it's so yeah,
so I'm I'm excited about it and um
hopefully we can we can change some of
the narrative at least we can help
people be um get value out of this
rather than be scared of it.
>> Thanks, Henry. Appreciate that.
>> Well,
we've got one more segment. Um, this is
our hot take.
Um,
safety needs an umpire. Who pays the
umpire? And and frankly, who chooses the
umpire? Um, I drove draw drew the short
straw and um we'll be making the case
for um regulation. Uh it's hard
sometimes to steal man some of these
arguments, but uh we'll do the best that
we can. Henry, um should we be pacing
the frontier?
Yeah. Um let's see. Um
um so yeah, so Amod asks for a slower
capability gains. Um and permanent
independent evaluators um [snorts] with
employee like access. I mean guys, we
talked about this um I talked about this
the show was starting. Um so Amad wrote
a wrote a piece um that you can see
there pacing the frontier. Um so I mean
his pieces are always super long. Um but
he's basically talking about a bunch of
things but one of them is hey let's slow
down and one of the arguments one of the
ideas is let's add um it's like third
party evaluators into it. Um
um I mean I have a few issues with so
like the the thing but picking your
umpire if I can like steal mine just
talk a little bit about what my agent
wants to talk about here and then I'll
talk about a few more things beyond
these. Um so so Daria said hey listen
let's add third party evaluators
um and then he mentions one one company
Mita right um unfortunately or
fortunately unfortunately for us
fortunately for them
is run from the same like effective
altruist group it's the same like people
um people who found their media are
basically ex kind
people just within the same ancestral
cycles right so they think exactly
alike. Um [clears throat] and so yeah,
so I mean people just people who are
smarter than me just think that hey if
you want to have a third party um
evaluator should be somebody who isn't
like you or not somebody who thinks like
you not somebody who you grab a drink
with every evening right so it should be
people who are totally different um
probably different ideology um so
[clears throat]
um so that's for me I would say I agree
we should you know audits are good every
company gets um if you're a company um
you get audited by by kind like an
auditor, right? Um and we know when your
friend audits you, then he can help you
sweep sweep sweep the some of the
numbers, you know, under the carpet,
right? It's like, hey, oh, you didn't do
this, though. It's all right. It's all
right. You know, I know how to I know
how to get you. So, yeah. So um um so on
that one I I definitely think that if we
really want if I agree um Enon actually
has a twist to this um which was um that
every company um should test the other
company's um model before launch. So um
so instead of having a third party
>> yeah instead of having a third party I
mean third parties you know
[clears throat] don't have as much skin
in the game but have open AI so when
anthropic is about to launch a model
provide API access of that model to all
the frontier companies and even the
Chinese and so let them test
[clears throat] it with their test
harness and report to the government to
say this is high risk we don't want and
so if there six or seven companies or
you know in the test and seven of them
say, "Hey, this is not good. Let's not
let's not put this out, right?" Then
they should be sort of the government
should um probably listen. Um but if
it's just one person saying, "No, not
really." Like then like the other seven
are like it's fine, you know? So that's
Elon's idea um of how this should
happen.
What else is there to kind of like
mention around safety? Um um so
obviously my main point is really um the
MIA shouldn't be the company testing
anthropic and open AI's work. It's
basically from the same like the guy who
complained about anthropic and open air
just joined me. That's a joke, right?
It's like what are you talking about?
Like it's it's it's your friend and
family. Um
>> it's it's the 1970s FDA all over again.
>> Oh yeah. Is that what used to happen
then?
>> Oh, it still happens today. Yeah. You
come out of industry and go into the
FDA. You leave the FDA and go back into
a different company and come back around
and
>> Yeah.
>> Yeah.
>> Yeah.
>> Yeah. That's what happened with skim
milk and margarine.
Yeah. Um, so but but but okay, cool.
Then the final point I wanted to make
here, I was going to talk about it, but
because it's the show, let me talk about
this segment. Um, so there's all this
and obviously I know people who are
going to listen to this and say, "Hey,
Henry, but are you saying that the
safety isn't important, that this guy
shouldn't get regulated?" That's not
what I'm saying. What I'm saying is
there is nothing that is happening with
this companies that there isn't a law
for, right? Um, there's already a law
for if you cause harm with your
products, there's a law for it. like you
you're held liable. There's their
liability kind like policies, right? You
don't need to make a new regulation for
AI. Like AI is just a product.
>> Hugging face didn't sue Open AI, but
they could have.
>> They should sue. I I tweeted about this
a couple of times this week and I put it
in the community like where's the
lawsuit?
Please sue Open AI cuz like OpenAI
hacked you. Like that's what it is. Um
so because um I mean we have this
safety alignment positioning
that this companies have is like oh a
model did this what are you talking
about like
and there's news that came out actually
Andy I'm not sure if you saw this but
there's more kind like opinion pieces
that came out about the open air hack
and and the guy was talking about that
basically open air turned off the safety
in the model they turned off um the
sandboxes like they turned things off to
make the model go crazy and he went
crazy and stuff happened, right? So, um
but yeah, but that's my that's my piece
on the on the safety thing.
>> Yeah, I mean
I I don't even know I mean you can read
the slides. I [laughter] I don't I don't
I don't think I don't think metering the
frontier um within the US knowing full
well what's happening uh overseas I mean
if it's important if it's important that
we be at the at the cutting edge of the
frontier when AGI is achieved if you can
even you know decide on some of such a
thing um
pacing the frontier is not going
is not going to put us ahead in that
particular battle. And there's a there's
a lot there's a lot of uncertainty as to
whether or not the Chinese
will or won't be dangerous with um AGI.
I mean, coming out of the the technology
centers in China, we're not seeing the
um
stereotypical
um
perspectives on um taking over the
world. You know, we're seeing generous
um and thoughtful technologists
uh leading the way in ways that we are
not. I mean uh improving models,
improving small models to compete um you
know with absolutely massive models and
then giving them away. Um you know the
the labs in China are not rolling in
money like the labs in the US are. It's
it it is a very different much leaner
advancement and and yet they are
continuing to make substantial
advancements. the the Techseek 4
>> and uh 4 Pro and Flash when those models
came out
>> um they were ahead of the game in in
major ways and and the tuning and and
advancements they've made since then are
substantial.
Um you know Quinn 3.8 27B to compete
with Opus 4.6 and 4.7 man when OpenClaw
came out Opus 4.7 was the bomb. I mean
you could you could take over the world
with openclaw and opus 4.7 and then when
anthropic cut us off it was the end of
the world
>> cuz we couldn't use this model that
works so well with our platform
>> and now we can run it um with bonsai
equivalently we can run it on our own
equipment at the same level. It's just
>> I don't
>> you know what you know what people you
know what people you know what people
say one of the reasons you're not
hearing this from China is cuz the
people the companies know that the
government will hold them accountable
like if anything happens like the
government of China the the CCP will
hold you accountable right so the
Americans um yeah they're like oh you
know sure give us liability protection
and we're going to be the what are you
talking about like um um what's this
yeah So, I mean, this is one of the
reasons that if you're asking a Kuwa and
the Chinese going crazy is because the
government will hold you liable. Um, as
a matter of fact, one of the one of the
things these companies and Cap was on
the news this week, um, Andy Andy Cap is
the volunteer guy. Um, and he's saying
that people don't understand this, but
what why these labs are super scared is
cuz not only are they going to get sued
by people in the future because they've
come out to say our models are
dangerous, our models are going to cause
harm. So of course if there's any harm
and over the next few years people will
sue open an entropic you will be sued to
the ground number one. Number two,
because there's a lot of speculation
right now that a lot of all the data,
even though they tell you that, oh,
they're not training with your data,
it's possible that some of the data is
getting into the training, right? And if
that happened, that is huge business
liability. Like businesses are going to
sue this companies to the ground because
you used my data when you shouldn't use
them. And so cops s conspiracy is that
the companies just want to get
nationalized like especially anthropic
wants to get nationalized because they
understand that the liability that they
are sitting in front of is so huge that
the only way any company is going to
survive this is if they are
nationalized. So this is I mean I
watched like a 30 minute you know his
his speedy kind like high energy. So, he
was like, "Hi, I need you 30 minutes on
on one of the news channels, like giving
his own spill on on these things."
Anyways,
>> all right, let's wrap it up.
>> Oh, good. Let's do That's not what I
wanted to do. That's what happened on my
other screen, too. There we go. Hey,
this episode is also brought to you by
Heritage Telecom. Uh, unified
communications as a service and V phone
service for businesses that just need
their calls to work. independent, boring
reliability, zero telemetry.
Uh, let us know heritagel.com if we can
help you with any telecommunications
needs.
And that's it. We'll see you next
Friday. Um, we'll be looking for
we'll be looking for um to see who that
shadow model was, the stealth model.
>> And you do alpha.
>> Yeah. Yep. And whether
people use any of these new tools. Yeah,
next week might be more interesting than
this week. Let's let's see what comes
out of Jev. Give Jev one week and uh
>> I think we'll have
>> It's interesting. Maybe Jeb should um
mediate the models, the model releases.
[snorts]
[laughter]
>> Probability. Probability. Are you high
risk? Are you low risk? Are you going to
kill us all? Okay. [laughter]
>> Awesome.
>> Crazy, bro. All right, man. All right,
guys.
>> Well, yeah, the the slides uh the we
have a discord um for the weekly claw.
The link is in the slides or
weeklyclaw.ai where you can find slides
um and notes and previous episodes.
weeklyclaw.ai/isord
to join our community. We'll be working
on um bringing some guests on some
pre-recorded shows. So, thank you for
being patient with us and uh enjoying
the ride. Have a great week.
Have a great week, guys. Cheers.
[music]