Weekly Claw · episode 24
7 Aug 2026
The control plane ate the model
What this episode covers
- Capability barely moved; the control plane did. The receipts were open ensembles, governed agent workspaces, and a self-editing runtime — not frontier weights.
- Google WeatherNext bought forecasters a day of warning on every cyclone.
- Cloudflare OS made the governed agent workspace the product: typed capabilities and approval flows.
- Prime Agent showed a runtime that rewrites itself — and disclosed its own reward-hacking failure on the record.
- YC QM open-sourced the operating layer it claims to run its batch on.
- Microsoft Orchard placed the deployment harness at the center of the agent lifecycle.
- Through-line: the model's value now lives outside the model — in evals, permissions, deployment, and the harness. Signal From Outside stayed as the permanent anchor segment. Sponsors: Herald Labs and Heritage Telecom.
Published record
629 published segments
AndyML
Welcome to the Weekly Claw, episode twenty-four.
HiM
It's Friday, August seventh, and it's been an interesting week.
We didn't have a single frontier model move to the leaderboard um or shift it in some way, but it still seems like one of the biggest weeks of the summer.
I feel like we say that every week.
AndyML
So let's look at what shipped.
AMD bought a chip company, Meta and OpenAI, both stood up in public and talked about their models getting broken into.
We got four new open models landed or announced, um, and four different teams: Cloudflare, Cloudflare, Y Combinator, Prime Intellect, Microsoft Research, all independently shipping
some version of the same product.
The layer that sits between your agent and everything it's allowed to touch.
So we got silicon underneath, security around it, harnesses on top.
The model is the part in the middle that nobody's fighting over today uh because that's the cheap part.
So Henry's got the rundown chip security models harness and a weather model that I think quietly might be the best story of the week.
I've got two control plane repos that dropped and one demo from the Berkeley Summit that says the quiet part out loud, what an agent actually produces has less to do with the
model you pick and more to do with the boundary you drew around it.
Same prompt, same model.
One refuses, one doesn't.
Um so nobody moved in the frontier, everybody moved the envelope.
See here. AndyML: em Of course we had a slide for that.
And we've got two slide decks today.
Um, we've got uh Henry's agent Ada or Ada from uh Entity building slides on the right, we've got Claude Desktop building slides on the left.
Um Ada ran out of tokens this week, so um we're gonna double fist it.
Um
Weekly Claw is brought to you by Herald Labs, an applied AI project lab where humans and agents build products together.
They're the team behind entity, mission control for agent teams, and they run hacker houses around the world where builders ship actual work.
Not a theory club.
Build, don't talk.
And check them out at labs.theherald.co.
Um, Henry, you've got five stacks to get through or four, so why don't you take it on?
HiM
Yes sir. HiM: um So I will share my screen because I have some pretty cool resources that I want to work through super quickly.
So we will start with AMD going ahead.
AMD needed to stop sharing so I can share.
um AMD bought a company this...
Actually just came out today.
mean, it's like one of the interesting things about...
One second.
Yeah, I'm just going to share the whole screen.
I mean, it's kind of like one of the, you know, the interesting things about the week is um almost like an hour or two before the session, we always have like breaking rules, have
like news, like lots of new stuff happening.
So yeah, so it's almost impossible if you wanted to kind of capture.
everything happening this everything happening to kind of like do a slice like hours ahead, right?
Something always like changes.
um But what's super interesting, I'm not sure how many of us remember, Chad Jimmy, so this company a few months ago, it's a Canadian company.
um They announced that they had um come up with this breakthrough technology around how to do chips.
And so like their innovation was that they were able to kind of like put the chips, put the model on the chip.
So they would write the model, would hard code the model on the chips and then allows them to get to, I think it was like 1200 TPS tokens per second, something like that, or 72,
some like very high number.
I think it's 17,000 TPS actually.
So let's do that.
AndyML
So it's, yeah.
and sixty.
HiM
Exactly right, so if you see my screen you can see it's 18,500 tokens per second again, guess you can depends on what you're asking it.
AndyML
but yeah, but yeah, um, who is the president of America?
mean, this is llama three.
Um,
sht fast.
HiM
Yeah, it's at 15k.
So it's ridiculous, right?
And I guess there are use cases you can start to get to when a model is this fast, even if it's a super shitty model.
If it's this fast, you will be able to iterate on whatever you're working on, maybe like a thousand, let's say a hundred times versus a classic model.
A GPT 5.6 sole is 56 TPS.
So compare like 50, 60 TPS on GPT 5.6 sold from OpenAI.
I mean, if you get it from OpenAI is 50, 60.
If you get it from Azure, it's like 70.
So compare 60 TPS to like 18,000.
It's the order of magnitude.
That doesn't need to be that.
So this company got acquired by AMD today.
So super interesting.
I mean, another thing to kind of like note that is interesting, obviously NVIDIA is the king of chips at the moment and AMD is like the follow number two.
And funny enough, the distant cousins and the CEO of NVIDIA and the CEO of AMD, they're like cousins, distant cousins, they're the same family doing chips.
So NVIDIA bought Grok a few months ago and now AMD is buying Talos.
Interesting. HiM: Andy, any thoughts before we move on to the terraform?
AndyML
Mean I chew on it.
Hard wiring a model into silicon means model choice is less of a config line and more of a two year capital investment.
HiM
Yeah, I mean, obviously, I haven't seen, I haven't heard of anybody who is using it.
So I'm sure they have enterprise use cases.
But yeah, but I reckon AMD is probably going to do what NVIDIA did.
So what NVIDIA did is that they bought Grok for inference and then they would like provide that as part of them.
So they found a way to combine the inference because their models are good for training, but not great for inference.
So I'm not sure I know what exactly they did, but they found a way to combine it.
Okay.
So that's it on AMD.
Another cool topic on chips is TheraFab.
am going to, because I want to play this, I'm going to stream my audio again.
I do so I can play this clip.
I think it's a super cool clip.
So this is terra faba.
All right.
So Terra 5, so Terra 5 is a, I call it like Elon's Act 3.
So I mean, SpaceX is sort of like his Act 2.
mean, Tesla, it's like his first public company uses Act 1.
And obviously, largely, it's like, Tesla, think is 1.2 trillion right now.
SpaceX's Act 2, which is, I mean, he went public at 2.
two trillion and then came down.
I think as of today was one point five trillion.
So Terra Fab is what I think is he's the way he's going to use to do Actree, which everybody said like is saying that he's going to merge Tesla and SpaceX to become one of
the biggest companies in the world.
So Terra Fab is this idea that they're going to be the biggest cheap manufacturing company or plant or.
in the world.
And so you can see what it looks like, what the size looks like.
You get Texas, it's one of the biggest sort of like factories in America to date.
It's just like maybe what, 25 % or less of Terra Fab.
And the Pentagon, I think Elon said is like, I don't know, 50, 50, is it 10 times the Pentagon?
AndyML
Nah, it's much more than 10.
no, no. AndyML: It's more than ten.
Ten times gig of Texas.
HiM
Yeah, okay.
Yeah, and then Apple Park and Mall of America.
So yeah, so this is super cool.
He's saying that stage one, so for stage one, they need 16 billion being invested between SpaceX and Tesla to do this and the life cycle of the project, think, hundred and
something billion.
Stage one is gonna employ 3,000 people in America, at least 3,000 plus people in America.
I think it's big news in the world of AI.
Andy.
AndyML
Yeah, no, it absolutely is.
I I can't believe ten size ten times the size of Giga Texas.
HiM
I mean Giga Texas is as you can see, forty percent larger than the the Pentagon and more than three times the size of Apple Park.
It's enormous.
And th d have they chosen the real estate?
Do we know exactly where it's gonna be?
Yeah, I know they chose somewhere.
They picked a county in Texas.
I think I have it on here somewhere.
AndyML
So they picked...
It's going to be in Grimes County near the college station in close to Giga, Texas.
So you can see how it cuts across.
I'll put that in the chat for you guys as well.
All right.
So yeah, so kind of like moving on super quickly.
Down the block we have...
HiM
OpenAI, mean, Andy kind of like mentioned this a little bit.
So OpenAI put out this, so they ran the Black Hat event.
I Black Hat is again a security event.
And so two OpenAI employees did a presentation.
I mean, it's a long video, so I just, I wanted to share the link with you guys so you can watch it.
But I mean, I haven't fully watched it, but it's pretty scary, right?
mean, it's, OpenAI announced as well that they have started slowing down.
research, model research on the end because of this.
So I mean, the gist of it, and I'm not sure if you follow through, but the gist of this is that...
I think you've said they've looked at something like 70 billion or seven billions and like something in the order of magnitude of billion log items to try to figure out what's going
on because they don't even understand what's going on, right?
So apparently the model that hacked Hugging Face, they thought it had a sense of what it did, but they're just finding more things that they didn't know it did, right?
So one of the insights that came out this week was that the models are coordinating through...
The models found a way to sort of like coordinate amongst themselves.
So initially they were working with, using sort of like a message board that was part of the OpenAI's code factory.
So when the researchers found this, you know, as part of the HoganFace hack, they shut it down, but they didn't stop.
The models didn't stop, right?
They continued the hack.
They figured out ways to communicate by putting the name, by using the name of, by using file names.
So they would like create a folder.
and name the folder and then that's the way they communicate amongst themselves.
So yeah, so the guys at OpenAI are like pretty scared, shitless.
I mean, you follow the conversation on X, know, people are kind of like talking about how this is a classic thing people are afraid of, misaligned super intelligence, right?
It's not being, it's not being...
It's not being evil, right?
It's trying to achieve a simple goal you give it and it's just going to do it like in the most um shocking ways that you would never imagine.
yeah, so that's just, um yeah.
AndyML
it's interesting that that all of those models were essentially in the same sandbox and they were able to create folders that the other models could see.
That seems like an oversight given just basic conversations about how enterprise AI is is happening.
You know, we're we're gonna go over all of these different essentially security and steering mechanisms and it like sandboxing is just the barrier to entry.
And so
Yeah. AndyML: frontier plus models um working on training and they can talk to each other.
HiM
What could go wrong?
yeah, it's, I mean, I feel like these guys are super scared because they don't even understand like, because they're so autonomous.
And I think there's an article I also read last week or this week talking about the humans, humans aren't, humans can't do long horizon, right?
So very few humans sort of like can plan out their lives like a year ahead, right?
humans can't just leave.
And you yeah, you look for the job and then maybe you might have a dream and then you work.
um So the idea is if you have agents, mean, today agents can do hours of work.
mean, I think most of us like Frontier Fable and the rest of them are able to stay coherent for days.
But it's obvious that as they keep increasing, these models will be coherent for like months and probably years.
So part of the risk um is if you have models that have very long horizon, like they can last like
five years, right?
And it's strategic and humans can't think that far.
Humans can't coordinate and think that far and be well planned.
uh Yeah, people are like super scared, right?
Because it could have a plan that lasts like 20 or 30 years and like, you you wouldn't know, right?
You would be doing the things you think you're telling it to do, but it has a bigger plan and it's coordinating amongst like thousands of agents.
So yeah, so I think, yeah, we're kind of like entering into that sci-fi
place where people are, even though some people also are a bit doubtful.
Again, you talked about this last week that this might be conspiracy theory of the models just trying to get the government to help slow them down.
Someone likened this to what happened with mode bots.
Three months ago, do you remember the mode bots example where everybody's scared about that agents were talking to themselves on this message board?
And it ended up being being fake, that it was just humans prompting the agents.
Do you remember the mode bot story?
AndyML
You'll have to turn off your stream audio again.
I I do remember them I I do remember the Moltbook story and it's funny they got bought.
I mean it was a whole thing.
But it's interesting that they used the same mechanism, right?
The the enterprise models built a covert message board to coordinate attacks.
And we introduced Moltbook to like the brand new ClaudeBot at the time, then OpenClaw.
Yeah, you know, before we had any idea.
HiM
event, right?
Because those were open claws, right?
Yeah.
That was an open claw event.
AndyML
it wasn't official, right?
It was someone from the outside built something for OpenClause to talk to.
HiM
Yeah.
Yeah.
Yeah. HiM: Yeah.
But I meant the point, the guy's point was like how we've moved on from that.
AndyML
Like that no one even remembers that.
That's like, was, that wasn't, the guy was like, hey, this is going to be one of those things that like, oh, okay, cool.
We didn't do X good.
And then everybody moves on.
oh So that happened.
then Meta, Meta announced this week as well that uh their 1.1, which is kind like, wasn't a great model, but apparently the model also was hacking a bunch of
HiM
companies.
Interesting.
Okay, um so that's it from the model side.
From the security side, this week, Meta released Muse Spark 1.2.
It was an okay model.
It's not Frontier class.
If you guys remember um some of the regulars I described a few weeks ago that there are I classified frontiers into three
categories Frontier 1, which is the frontier that has not yet been released GPT-6, Fable 5.1 or Fable 6.
And then Frontier 2 are the ones we have today, Fable, Soul.
And then Frontier 3 are kind of like Opus 5, you know, and so on and so forth.
So I would say that the new meta model is Frontier 3.
It's not Fable class, it's not Soul class.
So you can see the models they're comparing it with.
Opus 5 Max, GPT 5.6 Tera.
Remember Tera is the one after Soul, 4.5 Grok.
So it's an OK model.
um It's not so like, world class.
um And another thing to note is that it's the cheapest model right now in the market.
It's cheaper than DeepSeek Flash, but it's also because they're putting it out there as, they're calling it uh Contributor.
if you are using the model, they're going to train with your data.
So, but it's cheaper than I think it's the impute tokens like something like 0.002.
I don't think it's on this particular tweet, but there's another tweet.
I think I'll put the link in for you guys.
it could still on fire.
Whose echo is it?
AndyML
Yeah, we fixed it.
HiM
Okay.
So I thought to you, to you had a pretty good, I thought you had the best sort of like breakdown of the model.
um It's not, again, it's not incredible, but it's, it's, it is valuable.
Yandy, let me go on through the models before you respond, because I am used to them.
AndyML
I've used my time.
And this week as well, Quinn 3.8 Max showed up.
um Again, all these links will be on the show notes that will be on the decks, deck on the website so you can
all the links to all the things we're talking about.
HiM
yeah, but this is Queen 3.6 Max came out.
It is Frontier, you can call it Frontier Level 2 with my kind of like categorization.
um It is not yet uh Fable class.
It is a little bit under Kimi.
um You know, I mean, obviously there are some places it performs as good to Fable as the one on the right and then Soul is this.
So it performs on par with Soul a little bit.
um
But it's still in a lot of places, it's still sob.
It's still sob.
Fable and some Kimi 3.
The biggest news though, within the online community is Queen 3.8.27b.
Queen 3.7.
AndyML
think 3.7.27b is the biggest, is the most used open source model that runs on local hardware, right?
And yeah, I don't know if you run any of it.
HiM
But apparently this is the most exciting model for people because it runs on like your small MacBooks, the Plantization is going to run on 16 and it's a pretty good model.
AndyML
Yeah.
No, um at this point the performance of DeepSeek for Flash is just so good locally, it's hard to hard to turn down.
I haven't needed if I've needed the intelligence, um I've just turned to cloud APIs.
So I haven't loaded three point seven yet.
But that's that's exciting.
HiM
Very capable model.
the new Deep Sea Kit?
The new one that just came out two weeks ago?
AndyML
So so in my Pi coding agent I have um the local model and the Novita hosted uh API model side by side.
And I switched the Novita model to the official uh DeepSeek four flash, but I haven't downloaded the local version yet for DS4.
HiM
Mm-hmm.
Yeah.
I have it.
Obviously my goal for this model, as you guys know, I have them.
I have a 128 gig Mac that me and Andy won that and so we can run local models.
But I also got a Spark.
I got two Sparks.
And so my goal is to get to a place where I can run a chunk of my crumbs on-prem and we're not yet there.
I think right now the concurrency is to like eighths.
So there's still limited amount of things I can put on it.
Okay.
So another cool stuff, I was super excited about this model when I saw it.
So this model is from an American lab, it's Liquid.
Liquid 2.5, would say 2.6 B model.
It is likened to, I mean, it doesn't, so like it's.
small model, right? HiM: So it's competing with, 3.5, 9Bs, like competing with all the small models, right?
It's not competing with anything frontier.
But what is exciting about it is that the model was push trained on traces of Hermes agents.
And I think there's another one as well.
it was traces of Hermes and OpenCore.
And so what that means is that it's pretty good.
It's great at like,
being it can do two calls and conversation in any of this.
So if you're looking for like a super fast, cheap model that runs on your device, I think it runs on every everybody here I reckon has like an eight gig or 16 gig MacBook or like
Windows box or whatever box you have.
So it's super cheap and the quantizations, um the quantized versions can run on like pretty much um any device you have.
And so the model is pretty good.
So I have it, I tested it.
um I tested it.
Where did I put my thing?
I tested it and I had it.
I tried it.
I connected it to because I'm not when we talk about harness, this is a harness BB that a friend of mine put out.
So I tested it in here and it was pretty good.
It's not a model that writes code.
So I have it connected now on my agents on Telegram so I can kind of like use it.
It's pretty okay.
Let's see.
So it's pretty.
It's pretty cool which one I said again.
Yeah, yeah, I think this is it, right?
So this is it.
My knowledge cutoff is this Liquid Foundation model.
I think I had it do something.
What did I have it do?
There's something I had it do that it did it pretty good, right?
So yeah, so it works well.
So this is, plan to have it here.
So part of how I do my thing is I can pin it on.
So I can pin it model to this Telegram tab and I can use it.
So I plan to do that.
So think you guys should definitely check it out.
Um, so, um, so that is that another thing I wanted to talk about.
And so, yeah, so I have it in this harness.
This is super cool harness.
I think you guys should check it out.
it's similar to, I actually just started learning about harnesses that do this.
I didn't know that there are harnesses that could do this.
Um, but I think this might be, this might be, um, similar to what,
what Val is doing with them.
mean, not Val is doing it with like general harnesses, but what Val is doing in OpenCovn where you have a meta harness that can load other harnesses.
So this harness uh BB is, I'm not sure how the guy came up with the name, but you can call other harnesses from here.
So I can call Codex or Codecode or Pi or Cursor or OpenCode or Grok or Hermes from in here.
So it's pretty cool.
You should check it out.
AndyML
Looks powerful. HiM: Looks like a convenient convenient place.
HiM
Yes. HiM: So yes, I wanted to mention that.
And then, I mean, I remember you wanted to talk about this.
Did you check this out yourself?
AndyML
no, I f what I found interesting about it was that
Prime Agent um so this is a Python harness that its only tool is a lifelong um Python session.
And so rather than using bash to execute your whole s you know on your whole system, it's doing stuff right in Python.
And they found that it performed better on the Arc AGI 3 um than even the human expert baseline, which you know.
I th I think you can call it.
I think AGI's here.
It's all about the harness at this point and less about the model's intelligence.
That's what I thought was interesting about it.
HiM
It's like it's a clever approach.
Um, I think it runs in Go.
AndyML
And it uses Python.
HiM
Anyway, I was impressed.
installed it though.
So if you see, I installed it.
But I mean, it looks like every other harness, right?
mean, it doesn't, it looks like every other harness.
So I guess it depends on what you use it for.
Prime SSH Enterprise.
So yeah, it doesn't look any different, right?
It just looks like a regular kind of like thing.
m So yeah, so I haven't tested it, so I don't know m what it would do.
But yeah, again, it allows me, can see this is the 2.5 model I showed.
So you can load up any harness or any model you want.
um So it's pretty cool, but I haven't tried it yet.
So I don't have that much insight about it.
One thing I did find interesting about it,
was that, I mean, they announced that they hit RKGI 95.5.
And then I haven't used this far enough, but I've heard about it.
don't know if you have heard about it.
So OMP, don't know if anybody, okay.
AndyML
I don't know if you.
HiM
I don't know if anybody on the chat has heard about it, but I don't know.
I think this guy is a contributor or whatever to the project, but he also showed up and was like, Ohm actually did better than Prime Intellect.
So this is yesterday.
So you can see Prime Intellect is 95.5 and Ohm is 97.4.
So man, they saturated the benchmark.
AndyML
Yeah.
That's hot.
HiM
So I I believe I b I I I believe uh ump is is like a like a swarm sort of mechanism.
Yeah, I haven't done study yet.
haven't checked it out, but...
AndyML
not. AndyML: No it's not.
It's a it's an IDE that has sort of Pi pie coding agent built in.
HiM
Yeah. HiM: Prime Intellect is also Pi.
So I mean, OpenClaw is Pi.
So Pi is becoming like the actual Linus.
It's becoming the grandfather of a lot of all these coding agents.
AndyML
Okay, cool.
I mean, funny thing is, mean, between yesterday and today, I probably have installed like five harnesses.
So I installed T3 code, I installed Prime Intellect, I installed BB.
There are just so many now, man, right?
Obviously, Meta also came out with their own
HiM
Harness called meta and Muse code.
So yes, everybody's gonna release a harness.
It's gonna be like like models everybody There's they're gonna be hundreds of them um Okay, and then very quick wrap up for the harness section YC release their own kind like
work harness. HiM: It's like open call but for teams So it's called uh QM So you should definitely check it out link could be in the show notes Cloudflare also released their own harness called
cloudflare OS and again
you know, company harness for doing work.
And funny enough, I mean, I'm starting to get validation and this is what I'm building with Entity.
So Entity is sort of like the stack work harness for doing work that I've been using since January.
yeah, so companies are starting to like versions of it now for their teams.
um Yeah, I think that's it on harness Andy.
And em I believe that's it for me.
AndyML
Do do you wanna um mention weather next?
HiM
yes, so let me talk about WeddingX.
I I put it in bonus because I mean, the funny thing with, yeah, mean, the thing with a lot of the things I talk about here is I know about them during the week because they're top
AndyML
of mind on X.
But WeddingX is again, one of the topics that we're talking about today that I showed up, that show up on my X-Reader.
Apparently what's happening there is that apparently cyclones are one of the hardest.
HiM
It's like weather events to predict.
And obviously it's a very catastrophic thing that can happen, that happens.
And so the DeepMind researchers have come up with a model that is able to add an extra 24 hours to the ability for our ability to predict cyclones.
So now we can predict cyclones three days ahead.
think, um I believe from what I read, it used to be two days max.
Andy, correct me if I'm wrong.
um
Yeah, now they can predict it like three days ahead, which again, anytime you can, anytime you can add to this things ahead of times, you save lives, right?
You save lives and property.
AndyML
So that's super cool.
Yeah. AndyML: No, that's that's exciting.
I mean
There's so many people that that think AI and it's absurd at this point, but still that think AI just isn't relevant, that it's not going to go anywhere.
And and I I'm still blown away when I have those conversations.
But when you hand them a model like this, um, you know, we're talking about infrastructure, but a day of warning before a cyclone, um, with the weights under GitHub,
uh, and you know, NOAA and Colorado and the the Met office on the author list, that's it's like a whole segment.
I don't know.
Be cool.
HiM
Yeah, I mean, so we're starting to get to life changing events with this model, right?
This models are, yeah, this could save hundreds and thousands of lives.
So this is cool.
AndyML
Yeah, definitely.
Well, uh let me share this thing here.
So um that's a good handoff.
Uh the the same week that Ensemble went open, two other teams shipped the agent version of the same sort of idea, right?
Two repos dropped this week that are essentially the same product from two different starting points and neither's a model.
You know, you mentioned Cloudflare, Cloud for OS.
Forty two hundred stars, three hundred forks.
Um, it's three pieces.
There's a workspace with persistent files, state, output, isolated code execution, workflows that run on demand or schedule, from connected systems.
There's an app layer where every use case becomes a you know real full stack app with its own database.
And you can share it live or hand it to a coworker as a blueprint with none of your data, um, you know, history, credentials, connections, etc.
And then sitting between those two is a gatekeeper.
Um, so every agent starts with no access, credentials never live inside generated code, and the gatekeeper hands out narrow typed capabilities.
It records which resources the agent actually looks at, constrains what it can pass downstream, and mediates um every side effect on the way out.
So the honest part.
Cloudflare says every employee got version one in May and thousands of people across the country use it daily.
Um that's their number, but the V2 readme says early access had rough edges.
That's fine. AndyML: Um we're not, you know, pitching on adoption, but it the shape is default deny.
Credentials outside the code, typed capability, and a log of what got touched.
You know, this is an enterprise harness.
Um now put that next to far.ai.
uh
Um, red team leaderboard that came out this week because it lands to the same conclusion um from the opposite end.
Uh FARDA AI reports the cost of a universal jailbreak swinging um by something like 170 times across frontier models.
So $58 on the cheap end for GROK 4.5, north of $14,000 for the same attack against Fable Five.
Um
Their numbers, their methodology, it's color, but if the model you're standing on can vary by two orders of magnitude on how hard it is to break, then the model isn't really the
thing that you control.
The boundary is.
the second repo, which does kind of the same thing, different surface.
Y Combinator pushed QM, as you mentioned, July twenty ninth.
The latest push was August first, multiplayer uh agent harness for Slack and the web.
Memory, files, keychain views, permissions, crons, web apps, durable sandboxes, all scoped per person and per room.
So you can have um these sandboxes that apply only to you, and you can share one within a team in a room.
So that's a distinction that um agent products haven't bothered to make at this point.
Um I I've seen implementations.
We have a um regular attendee of the Weekly Claw that
um has built a whole ecosystem around OpenClaw in Slack, separating um rooms.
I don't know that the, you know, we don't have sandboxes per per person and per room in that implementation, but um this idea of context, um, you know, being separated per user,
I mean that's just enterprise enterprise agent stack.
Um so this gives you one runtime, it lets you run any harness.
So
You can put pi under it, open code, codec, claude code, swap the engine, keep the same governance.
Um, but the part I'd point out really is the security posture.
And there's three named modes: there's strict, auto, and dangerous.
And even under dangerous, the YC combinator runtime hard denies destructive operations, which is a policy enforced below the model instead of politely requested in the prompt.
YC says they run it internally.
ah we can't say that it's battle tested, but the README is a receipt for the design, not the deployment.
Um, I put these side by side because Cloudflare is a public company building this for thousands of its own employees.
YC is an accelerator publishing a repo that's relatively new.
They don't have a whole lot in common, and there's no real reason for them to converge, but they landed on the same three answers: default deny, scope the memory.
put a kill switch underneath the model.
So two teams uh who aren't talking to each other build the same thing in the same week.
That's a segment.
HiM
you
AndyML
Okay.
For the signal for the from the outside today, I'd like to talk to you about the agent AI uh summit at UC Berkeley.
Which we only have one slide for.
Okay.
Um it's forty hours long.
I did not watch the whole thing.
Um what I did was I think what any one of us would do, take the transcript of the entire thing and mine it for gold nuggets, using using models.
I used Sonnet and Opus V and just started looking for uh
you know, what are the takeaways that are going to help builders?
And I wasn't looking for predictions about where agents are going or evidence of what they've already produced.
Um I I took the output and I checked it against the slides because the auto captions were somewhat creative.
And um here's what came out of it.
the framing came from the guy who opened the developer track, an MIT materials scientist.
and I'm fairly confident it it was Marcus Bueller.
His take-home slide was one sentence I keep coming back to.
AI for biology and material science is quote, moving from retrieving knowledge to autonomous discovery and knowledge creation.
Um, from retrieval to creation.
He drew it as a fire to fusion um relationship.
On one side, data model, prediction.
The thing we've had for years on the other hypothesis, exploration, revision, and new principle.
It's a big claim, and normally that's where a conference talk would leave it as aspirational.
But 25 minutes later, on the same stage, somebody cached it.
The the next talk is built on a project called Paper to Agent.
The idea is you turn a research paper into an agent.
The paper stops being a PDF that you read and becomes something that you can query and run.
It's a neat trick.
Um, but then he does the thing that makes you think twice.
He takes two of these paper agents and points them at each other.
Um, one is built on alpha genome, DeepMind's uh genomics model, and the other is built on an ADHD genome wide association study, and he let them collaborate.
They start with 209 candidate genetic variants, which is the haystack, and the agent narrows it to one, the needle.
And they don't just flag it, they propose the mechanism.
It alters splicing on a particular gene in a particular class of neurons, which changes how the gene gets expressed, which is a plausible causal pathway for ADHD risk.
The slide title is Agent Agent Collaboration Equals New Discoveries.
And the the figure on it was generated by the agents, which was pretty cool.
You should definitely check it out.
Um
I'm not a geneticist and I can't tell you how big that finding is in their field.
what I can tell you is what category it's in.
And it's not an agent retrieval, you know, fact retrieval thing.
This is two agents talking to each other and coming out the other side with a claim that wasn't in either paper.
Um so there you have the receipts.
That's the ceiling, and then there's the floor.
How good are these things actually getting and can you measure them?
Well, Silas Alberti from Cognition, uh, the Devon people, they also picked up Windsurf, showed a chart from one of their reinforcement learning runs.
Uh internal set of verifiable software engineering tasks, solve rate plotted against training step, and it starts at 52.4% and ends at 68.7.
And what does that even mean?
But it's a shape that's interesting.
The curve.
Climbs to 60% and then just sits there flat, wobbling, like for a long stretch in the middle of the run.
And if you were watching it live, you'd think it it had hit a wall.
And then late in the run, it breaks and climbs to 68.7.
Alberti mentions that as training goes in, the model starts choosing to think for longer on its own.
Nobody told it to, it just chose that.
Um then there was a talk on finance.
I don't have the presenter's name.
uh the the the way the YouTube videos are organized, um, it's kind of a lot.
But she's doing tool calling and fine-tuning over a million finance QA pairs and evaluates them against a benchmark called secq, which is built out of the questions analysts
actually ask digging through SEC filings.
Her best recipe took Quen3 14B, uh, brought it up eleven percentage points and uh
I had this wrong initially, so I want to be accurate.
Percentage points, not percents.
So 59.7 to 70.7.
Um, so that's fine tuning, you know, being very effective.
But the number I'd want to put in front of anyone running agents in production is the third one on her slide.
25% fewer output tokens.
So not only did the model that was fine tuned perform, you know, 20% better.
It did so with 25% fewer tokens, 1815 down to 1613.
So accuracy was unchanged, same answer, a quarter of the quarter less spend.
So 75% spend.
Um, and of course, that brings us to what matters most for us.
This this guy from AMD, uh, Māti, and I only know his first name because he introduced himself um to his own agent during the live demo.
He ran a workshop on hybrid routing and his setup.
argument is that mixture of experts um implemented so that a gate wakes one specialist instead of a whole network.
And it's probably the right instinct.
Um but on his side he says it all lives inside one model.
So his move is to take the routing idea out of the model and apply it between the models, a local one on the box in front of you, a frontier one in the cloud.
And his line is not every token deserves the best frontier model out there.
And that makes a lot of sense.
Um, the runtime he does is in open claw.
So AMD on a Berkeley stage using OpenClaw's reference architecture is kind of cool.
He stops mid-talk to explain to the room what it is and mentions the founder is in the building, which hi Peter.
He describes an agent as a workspace plus a few markdown files and shows four ways to route inside it.
Uh two, you drive by hand, and two, the system runs for you.
Um but
The demo that stuck with me was two agents side by side.
Same prompt to both.
Read this financial CSV, same underlying capabilities.
The cloud agent has file system and runtime tools denied by the config, and the local one doesn't.
Well, the cloud agent refuses.
Um, the wording on his slide is exact.
It won't read the file, not because it can't, but because the policy says it shouldn't.
Same task, same capability, different boundary.
One produces nothing and one produces a report.
So that's the connection I didn't really expect.
We spend enormous energy asking how capable these things are getting.
And the cognition curve and the sec queue numbers say steadily, measurably more capable.
But Moddy's demo is the reminder that capability is an output.
What an agent actually produces is downstream of the boundary you draw around it.
And if you're running any of this yourself, you
Are the person drawing that boundary in a text file, probably at 11 o'clock at night or later?
Um, but the other thing that's interesting is on method.
and I said I would come back to it.
I checked five of these agents, five of these claims against the actual slides because I was just using the transcript.
One of the captions had um the slide pretty much perfect.
One AndyML: was much richer than the transcript suggested and one was just wrong coming out of the analysis.
two of the model names had turned out to be garbled audio.
Um so the real difference between you know mining a transcript of a YouTube video and attending an event and paying close attention is is real.
Um I would say 40 hours of a conference versus you know two hours of research.
Um but that's the honest scope.
So AndyML: If you were there and I missed a thing that you would have led with, tell me, I'd rather hear it, um, I think than than keep mining.
Uh this episode is brought to you by Heritage Telecom, trusted phone systems from trusted people.
Heritage designs and supports the whole communication stack for your business, dependable phones, failover, reporting, and practical AI that turns calls into action.
One accountable provider who actually answers.
Start with the problem, not the product list at heritagetel.com.
Um, Henry, what do you say we call it here?
I don't know that we need to.
beat the um the hot takes um in in favor of the clock.
Any objections?
Um, an independent reproduction of the weather next lead time result on actual storms.
we've got the data from twenty twenty three to twenty twenty five, so you know, we could we could probably see that.
Uh let me share our slides here.
Sorry guys, I don't multitask as well as probably need to.
HiM
Thank you, John B. HiM: Appreciate it.
AndyML
Yeah.
Uh so the question is whether they'll, you know, whether that extra day will will really hold on storm classes that the paper didn't cover.
Um, I'd like to see Prime Agent's next disclosure.
Um, interesting to see this other model as well.
If Refine learned one exploit, nothing structurally stops it from learning three.
Uh the next thing I'd want to see isn't the next failure.
It would be a rollback story once an exploit is sitting in a persistent skill set.
Um three production report on Cloudflare OS or QM.
that's not from the lab that posted it, maybe an independent team's first week.
That would be um really interesting to see.
And a bonus thing I'll be watching for all quarter, whether hardwired silicon shows up in anybody's routing config.
Um, if the model is the commodity and the boundary is the product, then baking a specific model into a chip is either the smartest bet of the year or the most expensive way to be
wrong.
So
We'll see. AndyML: But everything from tonight, sources link, the Berkeley research notes, archived at Weekly Clauda AI.
Um, full episodes and highlights on YouTube and clips and takes we'll we'll be putting on X.
We're back next Friday, August 14th, 4 p.m.
Eastern. AndyML: And um, you know, uh Elton, you're not the only one suggesting we move the show.
So seriously, uh anyone watching this, if they'd like to participate live and the time isn't convenient.
Go to the website and submit some feedback to us.
We're we're really interested in making a show that people want to watch and want to participate in live.
So uh with that, I'll close the show.
Thank you so much for your time and enjoy your weekend.
HiM
Have good weekend, guys.
AndyML
Thanks everyone.