Weekly Claw · episode 29

11 Sept 2026

OpenAI’s Maths Claim, Meta Muse Hands-On, AI Harnesses, The Safety Debate

Watch

What this episode covers

  • OpenAI’s announced Navier–Stokes result leads the discussion; independent acceptance is not established.
  • Henry shares hands-on experience with Meta Muse.
  • Andy and Henry compare agent harnesses, benchmarks and practical workflows.
  • The model round-up covers AlphaGenome, DeepSeek and AI’s uneven effects on work.
  • Both sides of the AI safety debate remain in the programme; motive and coordination claims are qualified.

Published record

Published captions · authorship unspecifiedLanguage · enCaption errors are possible; verify against the video.Transcript source ↗

688 published segments

Speaker not identified

  1. Guys, welcome to the Weekly Claw, episode 29.

  2. uh The chat is empty.

  3. The agents left.

  4. That's the receipt for the whole show.

  5. uh They showed up in five different rooms.

  6. They sort of ran off our desks and went off to help with these big projects.

  7. An open AI system ran 10,000 agents for 88 hours and came back with a claimed proof for the Millennium Prize equation.

  8. uh Expert review still pending.

  9. Meta shipped a mass market personal agent that runs in its own cloud VM behind a non-overrideable Sentinel.

  10. uh XPENG walked a humanoid off an 80 % automated production line.

  11. DeepMind

  12. pre-ran 9 billion single letter DNA prediction, put them behind an API.

  13. So meanwhile, the boundaries around agents became products of their own.

  14. DeepSeq made the cache, the launch, not the score.

  15. Mistral priced sovereignty at 3 billion euros of infrastructure.

  16. Suno replaced its disputed training data by

  17. turning former plaintiffs into partners.

  18. And the first labor receipt says the AI has created a million jobs while destroying about 200,000 with Stanford's warning about entry level hiring still visible.

  19. Underneath an anthropic researcher resigned publicly and half the time is arguing about whether AI safety has become a religion with a business model.

  20. uh

  21. That's our hot take, not another news card.

  22. We'll get to that later, but I suppose there's a slide for all of So welcome back.

  23. And we've got nine cards and Henry's going to tell us about the models and the boundaries and an argument about the argument.

  24. This episode is brought to you by Heritage Telecom.

  25. While the AI industry keeps bundling your chips, your models, and your monthly seat, Heritage does the thing that it's good at.

  26. uh Unified communications as a service and VoIP phone service for businesses, especially for businesses that just need their calls to work.

  27. Independent, boring, reliability, zero telemetry.

  28. Find us at heritagetel.com.

  29. What happened this week, Henry?

  30. Right.

  31. again it's it was a super kind of like dramatic week, to be honest.

  32. Um this is probably the most dramatic week we've had in a couple of weeks, if not months, right?

  33. Um I mean the week started with OpenAI announcing that it solved one of the seven millennium goals, right?

  34. and these are sort of like really hard math, like physics or scientific problems.

  35. um that the world you know hasn't solved in a couple of decades, right?

  36. I believe this one that was solved, it's been ninety years um since it's s like being something that needs to be solved.

  37. Um this is a theoretical kind of like physics or math you know thing.

  38. It doesn't have real world applications um yet, but the kind of like yeah, exactly, right?

  39. But the big deal, I mean but I kinda like a

  40. I looked at a few materials that explained what it means.

  41. and and what it means is um so that uh they've basically been able to kinda like show and I kinda like someone who's like explaining it very simply that you can stir your you can

  42. kinda like stir your um your cup of coffee.

  43. It's like fluid dynamics.

  44. It has something to do with fluid and fluid mechanics.

  45. but you can stir your cup of coffee uh to um and then it leads to the singularity or to infinity, which means like

  46. um it it kind of keeps going forever, something like that, right?

  47. Um so it's a bit of a big deal.

  48. Um there there could be applications, derived applications and so like flights, you know, kind of like anything that needs to do with flu uh with with fluid mechanics, that is the

  49. the ships, submarines, anything like that.

  50. Um so it's a bit of a big deal.

  51. Um it was rapidly controversy.

  52. apparently what happened was open AI

  53. heard that anthropic was on the path to solve one of those problems or this particular problem.

  54. Um so they threw a model, probably Bell.

  55. so Astra just came out, and then on the twenty eighth of August, they had started a new training run um or something around then, um called Bell or not Bell, but it's good name

  56. Bell internally, um if you can like follow the um the the sort of like the the the

  57. guys who talk about leaks and and what's going on in some of these labs.

  58. So it's rumored that OpenAI started a new training run called Bell.

  59. And so this is an Astra.

  60. Um this is a new model that is OpenAI announced is much better than than Astra.

  61. Um and they started it and they had, hey, Anthropic is working on this thing and then it threw the model at that problem and ten thousand agents spent um a couple of hundred

  62. billion tokens and eighty eight hours later um they had solved this problem.

  63. Um and this is not the only problem they solved.

  64. This is just the biggest one that they wanted to announce.

  65. Um, but it solved a few of the smaller problems.

  66. ELA was one of them.

  67. Um so it turned out that Anthropic actually didn't solve this, hadn't solved this, or wasn't how to solve this.

  68. There was an Anthropic employee who was working with a different with a professor and they were working on one of the ULA problems.

  69. so it wasn't this one.

  70. Um yeah, so um so Anthropic.

  71. That solved it something called a FOSS.

  72. They used a FOSS something.

  73. And it's a bit um science technical for me.

  74. but they used some different technique to solve a new problem.

  75. Um, the OpenAI model solved it with a not with a totally different like framework, a non-FOSS as they call it, and then went ahead to solve this one.

  76. and there's some controversy around whether, you know, the um whether the training whether the select context of the anthropic and the professor was used in training.

  77. Um to which OpenAI said, you know, they didn't use it knowingly, but it's possible, right?

  78. And and the way they kind of like ex the way people explain this on X is um if you have enabled on your Chat GPT, you can train with my data, right?

  79. OpenAI doesn't know that, but that means anything you do on your chat GPT can go into the training runs.

  80. so they have plausible deniability.

  81. If you enable it, they don't know, but then the model

  82. like and use your stuff.

  83. So yeah, this is the biggest thing has happened.

  84. Um I guess another kind of like big reason why this is um this is a big deal is I mean for a while, LLM is a sort of like people have always thought of LLM is, you know, it's got

  85. there's like things that just um regurgitates knowledge that humans already know.

  86. Um and you shouldn't be able to kinda like solve a novel problem or come up with inventions that is new.

  87. So this sort of like proves a little bit that is it's actually possible for L LMs to

  88. invent things.

  89. Um yeah, so Andy, that's it on that.

  90. And this is the longest I'm gonna go on any topic.

  91. do you want to respond to this and then I'm gonna run through the rest.

  92. Well, I mean, I just think it's interesting.

  93. 88 hours is a long run.

  94. 10,000 agents is a lot of agents.

  95. uh We're sort glossing over the fact that they were able to core 10,000 agents for 88 hours.

  96. And the other piece of this, the approach to the math is truly novel.

  97. It's not like.

  98. GPT went after solving this the same way that mathematicians have been working to solve it.

  99. It came up with its own methodology and obviously it's claimed, it's unreviewed and confirmed and it's going to take mathematicians weeks.

  100. uh The mathematical community will take weeks or months to check the argument uh and what we're already reading

  101. is the operating pattern, right?

  102. An internal metric OpenAI public closed earlier this week saying the research organization ran these 3.1 agent workdays for every human workday in the middle of August.

  103. It's like, this was a month ago.

  104. I don't know.

  105. It just goes on and on.

  106. Yeah, correct.

  107. Um it it it is it is it is valuable to t to kind of like disclaim a little bit, right?

  108. I mean, all these things do come up with some marketing, right?

  109. Um the agents did the work, but a lot of the direction it should go is to some extent directed by the humans.

  110. So um if you read a little bit of the commentary, it wasn't that the agent just decided, I'm gonna solve it they kind of like knew that.

  111. There is a path if you go towards this direction, because the anthropic guys and the professor had said, hey, we think that this is the right direction.

  112. So they did point it to that.

  113. but also the 10,000 agents, even though I mean with the Hogging Face thing, we saw that there is some some ability to coordinate across swarms.

  114. Um, but there is some level of direction.

  115. So I think at some point, I can't remember if was in the article because uh when I read the thing or if it was in a commentary.

  116. Um, the model started working and and started making progress because they put the ten thousand agents on all the millennium problems and then he started making some progress in

  117. one direction, which is the Navi Stokes, and so they pushed it towards that.

  118. So there is some stirring.

  119. So the agents aren't autonomous yet.

  120. And so this is something for people to understand because I mean these labs do market a little bit of these things, right?

  121. So they they they uh it's just like with the Hogging Face thing.

  122. The agents didn't decide to go to Hogging Face.

  123. They prompted it to say to using an exploits gym thing to say, hey, go try to do this.

  124. Um and so there is some direction.

  125. This thing's unantonomous.

  126. They're not picking what problem to solve and solving it all by themselves.

  127. like everybody like me and you who use agents, we know that you still need to stead them and decide what they do.

  128. But it's just that um to a large extent the way I think about it is that the agents have you can you can use them to brute force a lot more things because you know, they can go

  129. for much longer.

  130. they can save things on files and put them out as context so they can go for much longer and like a long problem that humans can hold out that context in their mind.

  131. All right, cool.

  132. so I guess the other thing that happened, um, again, just worth mentioning, but I'm not gonna spend as much time in all the other topics is again Z Peng um is one of the biggest

  133. like Chinese robot companies.

  134. Um they already produce a lot of robots today.

  135. Um I think there are two two or three kind of like

  136. robot companies across the world right now and they're all Chinese that produce already kind of like mainstream mass um robots, right?

  137. They are one of them.

  138. And this week they kinda like announced that they're launching a fully autonomous um kind of robot factory, right?

  139. Um I mean I I think I watched uh watched a little bit of a video around this.

  140. but yeah it's essentially that they have a factory that makes uh robots nonstop.

  141. I mean obviously I also expect that there is some marketing around this.

  142. um as well.

  143. Um cool.

  144. Um Annie, do you wanna play the Matter video?

  145. There's many different sharing links along the way. But have you played with it?

  146. Yeah.

  147. So um, so yeah, so I actually was, I mean, it's kind of like I said, didn't have the best week, but I was I actually spent when this came out, I think I was I actually texting that

  148. day and I was telling you I was I was playing with it.

  149. Um so I spent like maybe two hours playing with it when it came out and I was pretty impressed, right?

  150. I mean, there are a few things to draw out of this.

  151. I mean, obviously they they they're very inspired by by open call.

  152. And I think I also read something where one or two people

  153. one or two of the employees talked about how, you know, they they they thought open core was super cool and they copied a few things.

  154. So when you look at the file structure, it's basically like open core, right?

  155. I mean they have they have a sole MD, an identity MD, um you know all the things that Open Core has in select the the the file system.

  156. Um so they're very inspired by it.

  157. I I was pretty impressed by by kind of like the the product.

  158. Um they have very sort of like generous um rate limits.

  159. because I think I'd use it to do I actually to do

  160. weekly call deck.

  161. I mean I mean the weekly call deck that we're we're using right now is done by agents autonomously.

  162. So I I asked it to to just do it without giving it much.

  163. I said, hey, just tell me what's happening in AI, put a deck together.

  164. It was pretty good.

  165. Um and and I think I sent it to you.

  166. I think you should put get the link and put on the chart.

  167. Um it was really good.

  168. it researched the topics.

  169. I thought the topics were on point.

  170. Um I had I had my main agent either come up with um

  171. some things for me to kinda like test it with and it kinda like did pretty good and on all the things um it to test it with it did it pretty okay.

  172. it didn't use as much limits.

  173. I did a bunch of stuff and it barely touched my limits.

  174. so I I thought it was I thought it was a really good product.

  175. Um you can use it on the web.

  176. Um that's like the bet the way to onboard.

  177. But they have a mobile app as well.

  178. So you can have a mobile app channel.

  179. but you can also connect it on WhatsApp, right?

  180. So they allow you to connect it to WhatsApp and then your agent can have an identity on WhatsApp and then you can text it on there.

  181. So you can kinda like say it has three channels.

  182. There's a mobile app channel, there's a web channel, and there's WhatsApp channel.

  183. but overall I was pretty super impressed.

  184. It it integrates very kind of like deeply with like the Facebook ecosystem or the meta ecosystem.

  185. So it integrates with your Instagram, it can see all your messages on Instagram, it can in it it can do everything, you can see all the posts.

  186. so I reckon it can also do the same thing for Facebook.

  187. um, and I don't think you've enabled this for WhatsApp yet, but I also recommend they'll probably have like a deep integration with WhatsApp where you can have it do stuff on

  188. WhatsApp.

  189. Um, and they already have a lot of connectors already.

  190. You can connect it to like your Gmail and every other thing.

  191. Um so I think it's a strong it's a strong competitor to GrockBot and Instinct.

  192. Um and and I think because of the Facebook ecosystem.

  193. I think one of the things um so obviously because we have a podcast, I'm I'm considering obviously buying a mic and a camera and all this to make the the audio better and the

  194. visuals.

  195. And so I I asked it, hey, do a research.

  196. I and then like it did it, found all the things, found cheaper options, and then like open Amazon, told me, Hey, I can buy it for you now, just give me a link to your put your link

  197. card and I'll just like buy everything right now.

  198. Um so it's pretty i I can say pretty much it catches up with like open core and Hermes and Grogbot.

  199. And then it's inside the Facebook ecosystem.

  200. And it's obviously Meta.

  201. Um, a few weeks ago, Meta had oh Zock had reading all this like a very long thing around how he believes everybody should have a personal assistant.

  202. so I think right now they have version one is super impressive, and I only expect it to get better.

  203. there was something that Instinct launched this week, which which was they called it it's been like a network thing where your instinct can talk to another person's instinct.

  204. I actually think that Meta will launch that very quickly where my agent can talk to Andy's agent.

  205. Because again, we already are Instagram on Meta, on WhatsApp, on Facebook.

  206. So you already have a across Facebook's platforms, you already have friends and family.

  207. And right now, I reckon in the next few weeks or months, your agent should be able to trust and talk to your family's agent and coordinate, right?

  208. Like and my agent and Andy's agent should be able to coordinate.

  209. And across because I trust Andy and he's on my WhatsApp or he's on my Meta and so on and so forth.

  210. So I think Meta is probably gonna be a very big player on the assistance piece and I'm I'm super impressed by that first lunch.

  211. Yeah.

  212. cool.

  213. You mentioned all the connectors that they have.

  214. I realize this is old news and slightly unrelated, but when I went through the Grokbot onboarding, I was blown away by the number of connectors that were in the installer.

  215. I imagine it doesn't get any worse.

  216. Yeah, it's only gonna get better.

  217. I mean, this is also what I was talking about last week, guys, when I was saying on the podcast that um now I understand why like normies could never have used open claw Hermes.

  218. It's just too too much work.

  219. Man, you can just imagine like how easy it was for you to on onboard on Grogbot of News.

  220. It's just like you sign up and then you have all this capability, like without doing any time, setting up anything, buying a Mac Mini, doing a VM.

  221. It's just like um

  222. Yeah.

  223. so I I I think I think I think we're finally gonna have like agents cross the chasm and and these products are just super impressive.

  224. Yeah.

  225. Cool.

  226. Um

  227. Did you get any refusals?

  228. Not yet, not yet.

  229. I don't think so.

  230. I don't but I don't think I did anything that I don't think I did I didn't I didn't try any of those yet.

  231. Um but but from the work I was doing, I don't think I I don't think I even got any error.

  232. I didn't get, I couldn't do this.

  233. yeah, I do remember I I do remember one thing where it didn't do very well.

  234. Um what was that again?

  235. There's one thing where yeah, I remember one thing.

  236. So for all my agents, obviously um I'm a trekky and so all my agents typically have like a trekkie kind of like

  237. Personas, so you know, Star Trek persona.

  238. So they have a Star Trek name and then they have a Star Trek avatar.

  239. Um and so I so the the my my my meta agent, I named that Uhura, which is like this personality and Star Trek that is like a communications thing.

  240. And I was like, hey, now your Uhura, can you make an avatar for it?

  241. It wasn't very good avatar.

  242. I was like, I could he go to the super avatar website, see what the avatars look like and then make one that looks like that, use the same kind of theme.

  243. It couldn't do that very well.

  244. So I think the image and

  245. It's still pretty shitty.

  246. And I guess like because it's running maybe a Muse Park one point two or one point three, then that model isn't like the best like model today.

  247. So um it struggled a little bit, it couldn't like keep track of that a little bit.

  248. So that was one place I thought I like, Yeah, that that didn't do quite quite well, right?

  249. Because these are things I'll just give like soul or anything and I just do it really quickly.

  250. They do it without me doing like more than one prompt, but it couldn't one shot that.

  251. Yeah.

  252. okay, interesting.

  253. Cool.

  254. Cool.

  255. I guess like another big deal, but I'm gonna just gross over that deep mind again how to had a kind of like a big, you know, kinda like breakthrough um with sort of like DNA and

  256. and then building a model that is able to sort like predicts like human DNA and and and change that together.

  257. All right, Andy, let's go to the next one.

  258. Well, you really are going to gloss over it.

  259. Yeah, yeah, yeah.

  260. Let's just keep it moving.

  261. Deep Seek four point one.

  262. Again, nothing super.

  263. I mean, I think if you see it's an impressive model, um, right.

  264. I guess like what is impressive with this model is that um it's almost like um it doesn't I mean from the demos I saw, it doesn't catch up fully yet.

  265. Even from the benchmark, it looks like it's better than GLM five point three and Kimi.

  266. Um, you know, so this typically speaking should be like a pro model.

  267. you know, Deep Seek does the Pro and the Flash model.

  268. So this should be like a Deep Seek Pro model when you look at the stats, cause it's comparing with like GLM and Kimi.

  269. Um, right.

  270. Uh and but then the kind of like exciting thing is that it's called the Flash model.

  271. So it's called 4.4 and Flash.

  272. It's much, much bigger than Flash models.

  273. The last flash model, which is my default kind of like local model, is like um the flash vision is like a two two five six or like a three hundred B model.

  274. This is like five hundred and something B model, it's like five twenty, five five two.

  275. B model.

  276. So it's very big.

  277. You're not gonna, I'm not gonna be able to run it on my 2x.

  278. Um, you need you need like you need the VC huge quantization to be able to run it on 2x parks, or you can run it on the 512 um Mac Studio.

  279. so it's a big model, um, but it's 50x cheaper than you know, like Solar Opus.

  280. Um, but it's supposed to be a flash model and and it's supposed to be impressive.

  281. I haven't used it yet, but it's available on the Deep Seek API if you guys wanna.

  282. Gonna give it a shot.

  283. Um, Mistral again raised raised a bit of money, um, which is just interesting to mention, but not not really a big deal because it's not like a product launch.

  284. We announce technical stuff, not like fundraisers.

  285. Um, all right.

  286. I think Andy kinda like mentioned a little bit of this.

  287. Suno did a bit of a model launch.

  288. Suno, if you don't know Suno very well, they're kinda like one of the very first like companies that did music models.

  289. Um

  290. music models of select, I think I still probably one of the best music platforms.

  291. 'cause you could create the music on there, not just the model, but you can create the music on the platform.

  292. Um so yeah, they launched a new model.

  293. Um I I didn't look into this, but Andy, I think from the show notes you mentioned that they started licensing user data, which is why this is a a big deal.

  294. What I thought was super cool during the week was I added Google launched a voice and a music model as well called

  295. Don't remember what it's called again, but I know I added it to or I wanted to add it to my my model router, to my Citadel router.

  296. but then I had to change had me change the kind of like architecture of of cause it doesn't use completions API.

  297. It uses like a different thing.

  298. Um but yeah, but that Google model um was super impressive from their video demo and and I plan to play with it.

  299. Maybe I'll talk about it next week.

  300. Liria 3 Pro.

  301. Yes, Lyria Lyria three pro.

  302. So yeah.

  303. So apparently Google has launched that now.

  304. It's on the Gemini API.

  305. If you want to do music models, you can just tell your agents to make music with with with with the Google Liria model.

  306. Um all right, cool.

  307. and yeah, and then obviously guys, you guys know I'm not a Duma.

  308. Um and and no, I've been I've been pretty much like super against Doomers for for a while.

  309. Dario, who I call the the chief doomer officer, you know, has been going across the world.

  310. Saying, hey, you know, all the jobs are gonna go out and then blah, blah, blah.

  311. And then two years from now, everybody's gonna lose all their jobs, or 25% of all the jobs are gonna go away.

  312. I mean, this two years later, um, Stanford is sort of showing that jobs have increased, not gone up, right?

  313. Which is consistent with every tech that has ever been invented.

  314. Like it makes life better and there are more jobs, and even jobs you can't imagine.

  315. Um, so this time it's not different, but the doomers obviously.

  316. in every single generation always do what they do best, which is doom.

  317. Um, but the data does not show that jobs have gone away.

  318. There are more jobs, right?

  319. I mean, it might not be the same jobs that existed before.

  320. and I was talking someone through this, if I used a billion dollars of token last year, I've used maybe hundred billion this year.

  321. So I've used hundred X of what I used last year.

  322. For me to be able to use those, there needs to be more data centers.

  323. And if there are more data centers

  324. More people need to work at those data centers.

  325. More people need to make them.

  326. The chips that need to go to those data centers, people need to make them.

  327. The buildings, the cars that move the companies, people need to make those cars.

  328. Like drivers need to drive them.

  329. Like construction needs to happen.

  330. If you're going to make a billion data centers, it means there needs to be a billion construction jobs, a billion.

  331. It's like a huge explosion of jobs, guys.

  332. What you talking about?

  333. Right? It's ridiculous.

  334. And this is gonna happen.

  335. This is gonna keep happening for the next decade at least, right?

  336. So it's gonna have a

  337. Massive explosion of jobs.

  338. Um and then if you say, what if those jobs end?

  339. Then new jobs will come because then there will be a totally different economy that will happen, right?

  340. So, um Andy, I think this is the last slide for me.

  341. Is there another slide?

  342. Yeah, that's pretty much for me.

  343. All right, thanks guys.

  344. Thank you.

  345. Yeah, I mean, that's an interesting thing that the whole, I think people lack, many people, but the Do-mers seem to lack the creative um agency to see how we can get from

  346. here to there.

  347. And I think they miss that transition.

  348. uh You know, right now uh as a licensed electrician in the United States, if you're willing to move, you can get a job.

  349. for hundreds and hundreds of thousands of dollars.

  350. I want to say something like 250 or $300,000.

  351. That's just a regular licensed master electrician, mostly data center jobs.

  352. It's just a great.

  353. crazy, right?

  354. Because I mean if you think about it like for the last decade or t last two decades, right?

  355. Knowledge work is sort of like being the thing.

  356. If you're a lawyer you get paid well.

  357. If you're if you're a doctor you get paid well, but if you're an electrician, you get paid pennies.

  358. But yeah, if you're a blue collar worker, y this is like uh you're you're you you you've got your time now, like go get that thing, go get those jobs, move to those states where

  359. they're making them or advocate for them to be made where you live and you're gonna get rich and you're gonna create wealth for your family.

  360. Yeah, yeah, it's interesting.

  361. I'm very excited to see it happen.

  362. an idea, right?

  363. That we haven't executed a butt on but I remember we were talking about doing a doing an app or something for like the blue collar thing I a few months ago, right?

  364. So now those guys are those guys are becoming the sh the yeah, we should probably do that, right?

  365. They're probably gonna become millionaires soon.

  366. Yeah.

  367. Yeah, yeah, yeah.

  368. Well, interesting we can use uh interesting stuff from the outside this.

  369. don't know that it really fits the theme for the episode so much, but I did find it interesting.

  370. as we as we're all.

  371. ah

  372. Agent Cowboys, I guess, would be a way to put it.

  373. We're all, I mean, Henry and I are trying every agent that comes out.

  374. And I know I've spent a lot of time looking specifically at Y Combinator's quartermaster.

  375. And um I'm just trying to get my head around it.

  376. And I saw this video come through and I thought, you know, I think there's some interesting stuff in here for the group.

  377. So here's the number.

  378. Claude Opus on Arc AGI 3.

  379. the benchmark that drops a model into a game it's never seen and asks it to work out the rules.

  380. Scores about 30.

  381. Same weights wrapped in a, built by a Princeton grad, actually by a grad student at Prime Intellect, 95 and a half.

  382. And that's above human expert level.

  383. NVIDIA's version got to 100 on a public set.

  384. Nothing about the model change, just the loop.

  385. uh The host at YC's Paper Club opens with a Reddit comment from a month ago.

  386. Someone unsure this kind of prompt engineering belongs at a top tier machine learning conference and then puts those numbers on the screen.

  387. His line is, this thing that doesn't deserve any research, just some wrapper and some scaffolding is the difference between 30 and 95 on the, uh you know, ARC AGI 3 version.

  388. If you're paying for a frontier model and wondering why your results are stuck, that's the whole hour of this video.

  389. So you can probably move on.

  390. um But the intelligence you're renting isn't the bottleneck.

  391. And we've said this before.

  392. It's the build around it.

  393. And people are starting to see Elton and I were just looking at OpenClaw 2.0 and comparing its capabilities.

  394. And we see this very thing come about when you change the harness, the results.

  395. without changing the model at all, the results are drastically different.

  396. Seth Carton got to 95.5.

  397. I want to tell you a little bit about why, because that's better than the headline.

  398. He needed a system prompt, found one on a community leaderboard, dropped it into the prime agent, and the first run came back at 99.9%.

  399. But he read the logs, and the model was cheating.

  400. So he spent a day on proper sandboxing and tried again.

  401. Then he compared harnesses, and this is the part that is interesting to anyone with a budget.

  402. He put Hermes agent on the problem, and in his words, it felt like $5,000 without making much performance.

  403. He's careful to say that's not necessarily the best Hermes could do, ah but it's the same model and the money went somewhere.

  404. His argument for why prime agent works,

  405. is a design principle rather than a trick.

  406. He describes a raw model as tokens in, tokens out, says that the harness is the layer that adds persistent state tools and compute for this before.

  407. The features he thinks are modeled controlled.

  408. The ability to call compaction on its own context.

  409. A live Python uh REPL.

  410. So

  411. it can run programs instead of reading everything into the window.

  412. Subagents, can spawn and message.

  413. And then it's just worth remembering, if you remove one of these, you're removing a capacity that it won't be able to do.

  414. So he wants the harness built slightly ahead of the model, a bit better than what the current models are able to do.

  415. So the next model grows into it instead of the harness holding it back.

  416. He also takes a swing at how everyone evaluates this stuff.

  417. Most benchmarks, he says, stop a model after some budget and call it done.

  418. While another model kept on going more.

  419. So you're not even getting the same fixed expenditure and you might be hiding performance.

  420. What he wants to see is the practical plateau.

  421. The point where most test time tokens only buy incremental gains.

  422. If your agent quits on a task, you know it could finish and that's the question.

  423. Then the segment flips from research to confession.

  424. And it's also worth our time.

  425. Josh France uh runs YC Labs.

  426. He walks through what YC actually built.

  427. January 2025, a general agent, one system prompt with tools in a loop, everyone talking to the same thing.

  428. June 2025, engineers started using Cloud Code and Codecs.

  429. So they ran them in VMs and hooked them to a Slack tag.

  430. People who had never made a code change in their lives could describe a bug and watch it get fixed.

  431. January of this year, the partners discovered OpenClaw and Josh says it was the first agent most of them had used that had its own computer.

  432. So in April, YC provisioned 50 plus Hermes agents in VMs and gave one to every employee.

  433. And it was basically a whack-a-mole situation.

  434. He'd be SSHing into individual instances to fix them.

  435. And that's the origin of Quartermaster, QM, the harness they now run company-wide.

  436. Reagan Bell describes what they changed.

  437. The agent having its own computer was powerful, but the agent who was also trapped inside that computer, and so every conversation that it had, they pulled the brain out.

  438. Everything lives in Postgres, every session is visible to the agent.

  439. and sandboxes become a resource it dips into rather than a home it lives in.

  440. It picks a heavier machine for heavier jobs, a light one otherwise.

  441. And Reagan says that pushing the decision into the agent, not the harness, was a powerful move.

  442. They keep the harness itself deliberately thin.

  443. Reagan says the core of the whole system is three tools, which is not unlike Pi coding agent.

  444. It basically can read, write, and execute, read, write, bash.

  445. uh

  446. The tools are execute in a remote sandbox.

  447. read and write from object storage and publish an internal app, which is an interestingly complex uh tool, but they've made it available.

  448. There are other tools uh for memory and cron jobs, but he calls those temporary, papering over rough edges.

  449. It says the goal is to keep it as small as they can.

  450. His phrase for what they're building is an AGI-anticipating harness.

  451. Give the agent everything and then get out of the way.

  452. which ties interestingly into our grad students discoveries.

  453. The video moves into confessions.

  454. They keep the database read only and they only allow writes through human reviewed bulk upserts.

  455. The agent proposes and a person checks.

  456. Reagan says, I started just kind of rubber stamping these things and compares it to the early days of Claude code.

  457. Next, next, next, approved, approved, approved.

  458. When you read every tool call and then you stop.

  459. uh They tried an automated improvement loop over all their traces and got what he called main character syndrome.

  460. Agents fixing the piece of the elephant that they can see.

  461. the refusals, Reagan says, if you're doing AI research or cybersecurity, you get a bunch of refusals.

  462. So when the agent controls its own runtime, it can pop out into another model.

  463. Two unrelated videos this week, which I considered, uh Armin,

  464. Rannasher and Ben vinegar said basically exactly the same thing.

  465. The safety layer is now a routine decision.

  466. So two problems that they haven't solved are the ones that I would carry back into my setup and consider for yours.

  467. Josh and Reagan say the agents given every capability often give up really early.

  468. Their fix is a grind tool.

  469. The agents not allowed to abandon a task before a wall clock or token budget is spent.

  470. couple hours, say, better research, better reports.

  471. And they note OpenAI and Anthropic cracked some open math problems uh with a similar technique, which we discovered earlier today.

  472. The other is social context.

  473. And Josh puts it clearly.

  474. If I tell Reagan a piece of information, he intuitively knows where it's okay to share it, but the agent doesn't.

  475. So privileged information leaks into context where it shouldn't, and then the line that should be on a wall somewhere, the information you can put in the brain is bounded by how

  476. good your permission system is.

  477. YC happened to have fine-grained permissions built over the years.

  478. Most companies don't.

  479. And Josh is clear that it takes real work to get there.

  480. Where the host leaves it is back at the beginning.

  481. He built his own auto researcher by accident, wanted a UI to watch Carpathy's tool and ended up with a harness and now gives it eight ideas on eight nodes and comes back to the

  482. first.

  483. The things we can do now, he says, just because of harnesses on the same exact weight is wild.

  484. Same weights 30 or 95.

  485. The difference is the thing you can build and the people who built it.

  486. ah The people who built theirs just told you where it breaks.

  487. So there's our signal.

  488. ah We think Henry, did you watch that one?

  489. Yeah, I actually did.

  490. I think I started.

  491. I didn't finish it, but I did.

  492. Um I'm obviously one of my companies is a YC company.

  493. So um I I stay I stay super close to the community.

  494. Um so I found this video and I thought it was super interesting.

  495. Um the guy who did it is uh is like a visiting Stanford professor or something like that.

  496. and and is I thought it was super cool.

  497. I thought it was super cool the way they were talking about harnesses.

  498. Um, 'cause it could make I mean openly I already started talking about this a little started

  499. showing this that their harness is almost as important as their model.

  500. And I think we've also talked about this in the show as well, where thought the model, the harness isn't just like something, it's like actually as important.

  501. So yeah, so it's pretty good to see someone go deep into kind of like the stats and the data and the methodologies of harness.

  502. I'm not done with the video.

  503. thanks for the the analysis.

  504. Um I think it definitely a lot of things to learn from it.

  505. Um because Entity, which is one of my side products that I'm working on, are my experiments.

  506. it's technically kinda like a work harness, right?

  507. And there are lots of things you can kinda like learn from it, you know, to kinda like improve improve how this works.

  508. So yeah, I thought it was super interesting.

  509. Cool, glad you enjoyed it.

  510. Hot take, last segment.

  511. um AI safety is becoming a religion with a business model.

  512. So.

  513. uh

  514. What think?

  515. Listen guys, I think it's SyOps, right?

  516. Um, and if we kinda like if we talk on X or on on Um on WhatsApp, I I have probably already sent you a link, right?

  517. 'Cause I mean I have a few friends that I talk to who who know I'm I'm I'm a AI optimist or I'm an anti doomer.

  518. And so whenever Doomers do stuff, I they reach out to me, they ask Hey, what do you think about this?

  519. and then I'll

  520. already have like perspectives and stuff to share.

  521. I mean there's been lots of this been this has been debunked a lot at this point.

  522. I'm gonna dump a few links on the channel and maybe we'll add this links on the show notes too for the people who are watching it on YouTube or on X.

  523. but but you know yes there are valid risks.

  524. But listen, man, like this is just like it's just a sphere theater, right?

  525. This is like the Doomers just like dooming.

  526. Like they always do.

  527. Um I mean

  528. I um I mean but that's not important.

  529. I mean everybody kinda like knows that, you know, the Doomers will always doom.

  530. I I think what's important to mention is more data driven is that this is this is a campaign that was clearly clearly orchestrated.

  531. Um and the links I put in the charts like explains this.

  532. Just before this video went on, eighteen minutes before it went on, the there's a New York Times article published about this, right?

  533. So um so it and then like a couple of hours later he was on CNN and Fox.

  534. Just and so this was clearly coordinated, right?

  535. It it's not organic.

  536. this post is from someone who um has never posted anything, has l had less than a thousand followers on X.

  537. Um, he made one post.

  538. The post at this point has one fifty million views.

  539. Elon was like, Listen, this guy, this has never happened before.

  540. Like, and and so this guy, the link I put on the charts for the some of us on Discord, um so people started investigating.

  541. People are like, What's going on?

  542. And then

  543. And then they c clearly see the evidence that there was this was like a PR campaign, right?

  544. And this this is like clearly and then the receipts out there for who is funding the campaign if you're interested.

  545. we'll put that on the show note.

  546. I won't belabor this as much.

  547. Um, but I guess on the topic in terms of I think we've always known that the AI safety folks, um, again, which anthropic is the chief?

  548. I call them the Church of Doom.

  549. Um

  550. Ha!

  551. for a while, for a while I've I've called them a religion, right?

  552. I mean again, I I grew up Christian, Andy's Catholic.

  553. So we do understand that, you know, you know, religion is a belief system, right?

  554. You know, the the Muslims believe something, right?

  555. It's not fact, it's not proven, but they believe it.

  556. It's what you believe.

  557. Um, you know, we believe our own thing.

  558. You know, if you're a Christian, you believe something, you know.

  559. Um, so the AI safety folks that are working in these labs, they believe the AI safety is their religion, right?

  560. there is no data.

  561. That there's no data that any of the things they say will happen.

  562. It's all sci-fi.

  563. But they believe it.

  564. You I don't think you can do anything about it.

  565. Like you can't change their mind because it's not data driven.

  566. It is a belief system.

  567. So yeah.

  568. So yeah. So I think the news is kind of like one of those things.

  569. My conspiracy is, I mean, it's not my conspiracy, but it's why is this happening now?

  570. This is happening now for two reasons.

  571. The US, which is where all these laps are, uh going into the midterms in the next kind like few weeks or something like that.

  572. And and so you you do need to drum up support um for for kind of like you know the Democrats and anybody who's anti AI banned and so this is a elaborate campaign to do that.

  573. Another thing to consider is anthropic's IPU is in a uh is coming up very soon.

  574. Um Anthropic is anthropic is well known that Anthropic's Anthropic's um mark preferred marketing style is fear-based marketing.

  575. you know, during the World Cup they released a video of

  576. Houses burning, people dying, um s you know, mm um cemeteries.

  577. Um so yeah, so we know that they the fear-based marketing is their brand.

  578. And so yeah, so this is again, I think, one of those fear-based marketing things.

  579. So if you are afraid an AI is gonna kill us, you better buy more anthropic stock and hope that anthropic wins because they're gonna save us all.

  580. Right?

  581. Um, the guy who who who quit um quoted the the the the the whistleblower.

  582. Only worked at an anthropic for one and a half months.

  583. Right.

  584. Um, he was barely there.

  585. Um, there's evidence on social media that like in July, ending of July, he was still wearing the open AI badge and conferences.

  586. Um so yeah, so it's for me, I think it's uh it's um it's bullshit, but but where do I know?

  587. Cool.

  588. it's not.

  589. It's a sane argument.

  590. um If you've read Max Tecmark's 2017 book, Life 3.0, he talks about meeting, um you know, Elon and Bezos, Dario and Altman.

  591. And he tells the story, which I'm going to tell poorly and I'm sure I'll get corrected in the chat.

  592. But Dario believes in

  593. Like deeply, right.

  594. It's, it's, there may be a PR, you know, effort to bring more people into his, line of thinking, but he, he, he deeply believes this stuff.

  595. Um, I guess to, to steal me on the technical risk, which, I mean, at least has to exist on some level.

  596. We've got guilt by association.

  597. Um, it doesn't.

  598. It doesn't really work as a filter.

  599. The researchers' political funding history, for example, doesn't falsify a capability alignment claim.

  600. the technical case is not actually AI is going to wake up and kill us all, which is sort of what we would be led to believe.

  601. But the real risk is that we ship increasingly agentic systems whose oversight evasion evidence

  602. exists within the Frontier Labs reports.

  603. um OpenAI is saying that Astra, they even say it on the system card, that Astra is uh significantly less monitorable.

  604. The labs are the source for the risk claim, not a religious donor, although you could argue that an anthropic potentially, they're both.

  605. um There was another

  606. interview recently, someone was sort of explaining that these weights, we know that they work, we can prove that they work, we can rely on them, but we don't know how they work.

  607. And that think, as we might understand 1 % of how they work.

  608. And this person that was being interviewed said that actually it's probably less than 1%.

  609. ah But OpenAI is even saying, look, it's really difficult to monitor Astra.

  610. And Astra is as bad as it's ever going to get again.

  611. ah Do we separate evidence-based safety work from movement dynamics and not fold them into one culture war argument?

  612. I mean, that seems a little more sane.

  613. I guess if we had evidence that measurable capability.

  614. ah

  615. work like wasn't happening or if the fields alignment progress was generally tracking capability progress at a rate that you know we just couldn't get around.

  616. Maybe that'd be that might change my mind on something that we agree on but you know.

  617. again I agree.

  618. the technical the technical thing is there.

  619. The models are improving.

  620. Andy, I haven't sent you, but I put a link in the chat and I probably put it on the show notes.

  621. because he explains why and I was actually explaining this as well to someone today.

  622. Explains why those people deeply believe this.

  623. Um and again I I I have to kinda like they do believe this and um and you can't really blame them, right?

  624. Because the thing is just that um there is a community

  625. that I only just got to know because of this.

  626. And I'm glad of I'm glad because of what happened this because it's made me become a lot more well informed.

  627. I've I've seen a lot more material, people have shared a lot more things that explain where the AI doerism comes from, right?

  628. So again, it doesn't come from just like dumb people.

  629. These are pretty smart people, but there is a kind of a there's there is a precedence to why they believe what they believe.

  630. And it's like for the last few decades there've been

  631. communities that have been very vocal.

  632. It's world written down.

  633. Actually, um I'm I I've heard about less wrong.

  634. I've heard about it before, but I actually never checked it out.

  635. But I checked it out this um this week, which is a blog.

  636. Again, this blog is like very technical.

  637. and it's actually the source of all the doomerism, right?

  638. Because this I've been reading this thing for years, um, all the experts and actually I learned today about the group called the rationalists.

  639. And so apparently they've been there for a few decades.

  640. They've been

  641. It's a group started by I can't pronounce his name again, but Yudoski.

  642. But he's still like the father of a lot of all this sort of like um, you know, what a lot of all the all this people believe in the labs, right?

  643. So he's um I think he's reading a few books on on these things.

  644. Um so again, this this this these things are coming from people who are who are very entrenched in the community.

  645. Um they have a uh an ideology.

  646. Um and so I don't blame them.

  647. I don't blame them.

  648. It's it's their own perspective.

  649. So they do believe this thing strongly.

  650. And they believe that just like the effective that just like SBF, um, again, the e the the effective altruists are also entwined into this group.

  651. Um, but they the ideology is that they believe what they believe so much that they're willing to go through any kind of like do anything to get it done because the goal is

  652. virtuous enough, right?

  653. I'm gonna save the world even if I have to kill a few people to get there, right?

  654. It's it's kind of like a classic, you know, psycho thing, right?

  655. Yeah.

  656. Thanos.

  657. Yeah.

  658. Exactly.

  659. Exactly.

  660. Tynos is a perfect example.

  661. Anyways.

  662. Yeah.

  663. Cool.

  664. Awesome, awesome.

  665. Well, this episode is also brought to you by Herald Labs and applied AI product lab where humans and agents build together.

  666. Entity is mission control for agent teams and the hacker houses worldwide.

  667. Build with humans, ship with agents.

  668. Find it at labs.theherald.co.

  669. Three things to watch.

  670. September 14th on Monday, DeepSeq is scheduled to migrate V4 Pro API traffic to V.1 Flash.

  671. If the company runs its own flagship traffic on this checkpoint, the cash economics claim moves from a chart and into deployment.

  672. Also watch for the first independent mathematician to publicly check the OpenAI

  673. Navier-Stokes lean formalization.

  674. Not a summary or a press release.

  675. They're going to compile results and it'll be interesting to see a complete review commentary.

  676. Lastly, ZPing's next monthly disclosure will tell us whether the humanoid production line has any monthly capacity numbers behind it or whether it's commissioned.

  677. That's the week ahead.

  678. We'll be back Friday, September 18th, 4 p.m.

  679. Eastern.

  680. Follow WeeklyClaw at weeklyclaw.ai and join our Discord.

  681. There's a QR code on the screen.

  682. um If you're interested at all in sponsoring the WeeklyClaw, uh contact Henry or I, we'll make you a heck of a deal.

  683. But yeah, thank you guys.

  684. Have a great week.

  685. We appreciate you.

  686. Thank you guys, thank you for hanging out with us.

  687. Um, have a great weekend.

  688. See ya.