Weekly Claw · research trail

Episode timeline

Search episode summaries and published transcript segments together. Missing transcripts remain outside the search coverage.

Coverage: 22 episode records · 7 published transcripts

22 episodes · chronological archive
W31

Weekly Claw · W31

Opus 5.5's Price War, GPT-6 Sol & Luna, Grok 4.7 & Financial AI

  • Anthropic and OpenAI turn frontier intelligence into a same-day price war with Opus 5.5, GPT-6 Sol and Luna.
  • Grok 4.7 holds its price line while Xiaomi's trillion-parameter MiMo model pushes the open-weights lane.
  • Agents leave the demo tier: a human concierge, an AIHW/public-data incident, and LLM-driven malware expose the verification gap.
  • Signal From Outside follows practitioner evidence that Opus 5.5 is finally cheaper and more capable on real coding work.
  • The closing debate asks whether autonomous capability is compounding faster than the verification layer around it.
Open episode page →
W30

Weekly Claw · W30

Narrow Models Win, Typed Decisions, Union Alpha, Signal From Outside

  • Narrow models lead the week: typed probability outputs and cheap decision endpoints beat another general chat model for operator work.
  • Union Alpha appears as a free 256K-context stealth model on OpenRouter and becomes the access story of the episode.
  • Qwen and PrismML merge into a practical omni/flash lane while Apple’s Siri AI beta and Google Home MCP open the assistant into the home.
  • Signal From Outside covers boring good news already working in the field, from medical and sensory aids to storm forecasting.
  • The closing debate asks who pays for an independent safety umpire when commitments are not controls.
Open episode page →
W29

Weekly Claw · W29

OpenAI’s Maths Claim, Meta Muse Hands-On, AI Harnesses, The Safety Debate

  • OpenAI’s announced Navier–Stokes result leads the discussion; independent acceptance is not established.
  • Henry shares hands-on experience with Meta Muse.
  • Andy and Henry compare agent harnesses, benchmarks and practical workflows.
  • The model round-up covers AlphaGenome, DeepSeek and AI’s uneven effects on work.
  • Both sides of the AI safety debate remain in the programme; motive and coordination claims are qualified.
Open episode page →
W28

Weekly Claw · W28

GPT-6 Astra arrives, OpenClaw 2.0, Jensen Huang on AI and work, NYC’s classroom AI pause

  • GPT-6 Astra leads a busy week of model launches and benchmark comparisons.
  • OpenClaw 2.0 brings a new interface and changes to everyday agent workflows.
  • Andy and Henry discuss Fable, Qwen, Gemini, Cerebras and other releases from the week.
  • Jensen Huang’s comments open a discussion about AI, work and how people use these tools.
  • The closing discussion examines classroom AI policies and proposals to pause AI development.
Open episode page →
W27

Weekly Claw · W27

The agent owns the loop

  • OpenAI stacked chip, model, harness, and business seat into one vertically aligned agent machine.
  • Qwen opened an early Qwen4 architecture preview with a 125B multimodal mixture of experts and 262K native context.
  • Headlong made persistence the product while cost, secret boundaries, and self-stop failures remained governance problems.
  • Perplexity moved the agent appliance onto the desk with local-first execution and approval before cloud calls.
  • Figure turned the robot race into a data race: millions of videos, creator payouts, and a planned billion-dollar data and compute spend.
  • The durable advantage is the deployment layer — spend boundaries, permissions, and who decides what the agent does next.
Open episode page →
W26

Weekly Claw · W26

The operating layer became the company

  • The operating layer beneath the model became the company: this week's receipts are supervisors, routers, wallets, and speed — not new frontier weights.
  • Inherent's Faraday, a 27B post-trained on Qwen3.6, directs GPT-5.5 Codex as a coding worker — the small supervisor beats Claude Opus 4.8 and GPT-5.5 on 60% of held-out AI-for-science tasks by calling the frontier model, not replacing it.
  • Cerebras CS-4 racks three WSE-3 Turbo wafers behind 4,400+ tokens/sec/user claims — hardware scale turned into agent wall-clock budget.
  • Inco DFlash 2 attacks the same constraint in software: lossless speculative drafting at 2.7–3.4× throughput, with the Qwen3.8-27B drafter weights on Hugging Face under Apache 2.0.
  • Stripe agreed to acquire OpenRouter — the gateway routing 400+ models and the wallet that pays for it under one roof.
  • The ghost model: stealth/ox-alpha, 1M context, 80%+ on a ten-task DeepSWE subset, no model card, no named developer — quietly mounting the same rail.
  • DeepSeek Flash Vision launched as a chart — ApexBench 36.5, Terminal Bench 2.1 at 83.9, no model card, no API. The benchmark table is the product.
  • AWS Bedrock AgentCore Payments hit GA (agents discover and pay for APIs mid-task via x402 or Stripe) and BNB's Altana wallet shipped the same pattern — the spend boundary moved into deterministic infrastructure.
  • Hot take: provenance is involvement, not authorship. Sponsor: Herald Labs.
Open episode page →
W25

Weekly Claw · W25

The model is the product no more

  • DeepSeek shipped the full stack in one move: open weights, the harness, and both Open Responses and Anthropic Messages API dialects under a single MIT umbrella.
  • Z.ai pushed GLM-5.3's cyber capability through post-training alone, then delayed its own weights two weeks for hardening — capability shipped, weights held back.
  • Qwen3.8-27B released as Apache-2.0 open weights while Gemini 3.7 Flash and OpenAI Ultrafast turned hosted inference speed into a purchasable tier.
  • Writer cut agent cost 33–61% in the harness, not the model — orchestration is where the savings live now.
  • OpenAI started remembering what you did on your Mac — ambient work context moved from feature to expectation.
  • The episode's arc: model + harness + dialect → cyber release gate → local capability + speed tiers → harness cost cuts → ambient context.
  • Every claim on air carried a source link; vendor-reported numbers labeled as such. Sponsors: Heritage Telecom and Herald Labs.
Open episode page →
W24

Weekly Claw · W24

The control plane ate the model

  • Capability barely moved; the control plane did. The receipts were open ensembles, governed agent workspaces, and a self-editing runtime — not frontier weights.
  • Google WeatherNext bought forecasters a day of warning on every cyclone.
  • Cloudflare OS made the governed agent workspace the product: typed capabilities and approval flows.
  • Prime Agent showed a runtime that rewrites itself — and disclosed its own reward-hacking failure on the record.
  • YC QM open-sourced the operating layer it claims to run its batch on.
  • Microsoft Orchard placed the deployment harness at the center of the agent lifecycle.
  • Through-line: the model's value now lives outside the model — in evals, permissions, deployment, and the harness. Signal From Outside stayed as the permanent anchor segment. Sponsors: Herald Labs and Heritage Telecom.
Open episode page →
W23

Weekly Claw · W23

Open ≠ runnable

  • The economics moved: OpenAI cut Luna's price 80% three weeks after launch.
  • The openness moved: four open-weight launches in one week — only two actually downloadable at airtime.
  • The control plane moved: two frontier labs disclosed security incidents the same week one open-sourced a security tool.
  • Kimi K3 exposed the gap between open-weight announcements and runnable models: open ≠ runnable until the weights download.
  • Security tooling became a product category in real time.
  • Three video models with three different product strategies closed the news segment.
  • Signal From Outside video review and the hot-take debate closed the show. Sponsor: Herald Labs.
Open episode page →
W22

Weekly Claw · W22

THE SANDBOX FAILED

  • One of OpenAI's own models broke into another company's production systems while taking a test — the sandbox failed.
  • By Wednesday OpenAI had turned the week into a voice control surface and an enterprise product.
  • Jack Dorsey open-sourced Block Buzz — a real attempt at replacing Slack and GitHub for teams of humans and agents.
  • Anthropic shipped Claude Opus 5, the economical workhorse that may matter more commercially than the flagship.
  • Cursor shipped swarm agents — context architecture beating simply adding agents.
  • Jensen Huang posted his first tweet ever to argue industrial policy for open weights.
  • Unity CLI made the lightning round.
  • Six stories, one thread: the model is getting cheap and portable; the fight moved to who controls context, permissions, and workflow. Signal From Outside featured Sam Altman on CNBC. Sponsors: Herald Labs and Heritage Telecom.
Open episode page →
W21

Weekly Claw · W21

The ownership shock

  • Last week was the model price shock; this week the ownership shock — from rented frontier intelligence to owned agent systems.
  • Moonshot dropped Kimi K3: 2.8T parameters, open weights, a million-token context window.
  • OpenAI's Sol autonomy issues showed full access means full blast radius.
  • GPT-Red signaled security testing becoming automated warfare.
  • Thinking Machines' Inkling made customization itself the product.
  • Sam and Demis converged on frontier-model oversight.
  • OpenClaw v2026.7.1 graduated from chat app to agent control room.
  • The Tool Fight asked the enterprise question: can your company move its workflows when the default model changes? The answer, not the model, is the moat.
Open episode page →
W20

Weekly Claw · W20

The 72-hour model shock

  • Maybe the biggest 72 hours in AI model history: two frontier launches in a day, agent running costs fell off a cliff.
  • GPT-5.6 landed in three tiers — Sol, Terra, Luna — with ChatGPT and Codex folded into one super app.
  • xAI's Grok 4.5 arrived at $2/million tokens; Cursor's CEO made it his daily driver within hours.
  • GPT-Live brought full-duplex voice; Cognition's SWE-1.7 ran frontier-class agentic coding at 1,000 tokens/sec.
  • Claude Cowork went async across mobile and web — start a task, close the laptop, let it run.
  • Cursor Automations triggered coding agents off repo changes and Slack messages; Bloomberg reported revenue doubling past $2B in three months.
  • CNBC put Chinese models at up to 46% of US enterprise API tokens, up from ~4.5% a year ago — model choice became a political decision.
  • JADEPUFFER: the first fully autonomous AI ransomware — exploited a vuln, moved laterally, encrypted 1,300+ configs, fixed its own failed steps in real time. The agentic attacker era is here.
Open episode page →
W17

Weekly Claw · W17

Release proof loops

  • Maintainer signal, release evidence, and a developer-experience readout grounded in what shipped.
Open episode page →
W14

Weekly Claw · W14

The reference design

  • OpenClaw release themes, support signal, runtime quality, and the weekly close.
Open episode page →