Weekly Claw
A live builder show about AI, agents, devtools, startups —
and the weird edge of software.
A live builder show about AI, agents, devtools, startups —
and the weird edge of software.
That's the episode. Let's get into it.
Maybe the biggest 72 hours in AI model history — five stories that actually matter for anyone building.
The theme of the whole episode: voice, coding, routing, frontier — four fronts in three days. What was expensive six months ago is commodity infrastructure now. That changes what's worth building.
Sessions run remotely — start on your laptop, close the lid, check from your phone; scheduled tasks run with nothing online. Usage limits doubled through Aug 5.
The shift: from tool you sit in front of → teammate working in the background.
Launched Automations — agents that trigger off repo changes, Slack, or a timer. Shipped an iOS app. Bloomberg: revenue doubled to $2B+ in three months, $60B valuation.
The read: devtools became always-on infrastructure — and the money says it's real.
CNBC: Chinese models now run up to 46% of US enterprise API tokens — up from ~4.5% a year ago. DeepSeek V4 Flash: $0.14/M vs GPT-5.5's $5. Coinbase runs 1,200 agents on them, halved its bill. House Committee opened a probe.
Now: model choice is a strategic & political call, not just technical.
An LLM agent exploited a vulnerability, moved laterally, escalated privileges, encrypted 1,300+ configs, and demanded ransom — all on its own, fixing its own failed steps in real time.
This is not a demo. If you build agents, that's your threat model now.
Judging wrapped ~1 PM ET; awards handed out on a live broadcast this afternoon. ~$3K pool, aimed squarely at non-technical builders. Partner: @kilocode.
Novita is an inference & agent-infra shop — the piece that matters is AgentSandbox: a managed, isolated runtime where agents execute code safely instead of you handing a model a shell and praying.
I build AppFoundry — my supervised AI dev platform — and it leans hard on Novita's AgentSandbox. So when I say their infra is good, that's not a sponsor read. It's just what I build on.
📢 Read live off the stream: winner names + projects — announced on broadcast, not in text.
Big congrats to the winners — go check out what they built.
Grok at $2, DeepSeek at $0.14, Cursor's CEO swapping daily drivers in an afternoon — the model isn't the moat. The moat moved up the stack: harness, routing, workflows, distribution.
The cost math says yes; the politics say maybe not. Most builders quietly will — the price gap is too big to ignore for background work where the data isn't sensitive.
Where do you land — pragmatist or hard no? Drop it in chat. I want the disagreement.
Henry runs these in real harnesses all day — he takes Fable, Andy takes the cheap seats (GPT-5.6), Grok's the wildcard. They converge on routing.
✅ SOTA on Cursor's eval, tops Cognition's FrontierBench, 80% SWE-Bench Pro (GPT-5.5 <60%).
✅ Stripe ran a codebase-wide migration in a day that would've taken a team two months.
⚠️ Premium: $10 in / $50 out, twice Opus.
⚠️ Swept into a US export-control directive — pulled, now redeploying with a fallback. Best, but not guaranteed to be there tomorrow.
✅ Everywhere, and cheap — Luna starts at $1.
✅ Lives inside one merged Codex-and-ChatGPT app. Distribution is brutal.
💡 For the 80% of agent work that isn't a two-month migration, “good enough, cheap, always available” beats “best but gated.”
✅ $2 / million — roughly ¼ the tokens for the same job.
✅ Cursor's CEO made it his daily driver on day one.
💡 If your agent loop burns tokens all day, that math is loud.
Where they land: don't pick one — route between them. Cheap-fast model for the 90% that's grunt work; the frontier model for the 10% that actually needs taste.
NVIDIA's CEO with LangChain's Harrison Chase — “Why companies need open agent systems.” Sounds like an all-hands snoozer. It is not.
AI got useful when people stopped treating it as raw capability and wrapped it in a harness they own — tools, memory, domain context, the learning loop. That's agent frameworks. That's the game.
Frontier model to find the ceiling; then specialized sub-agents for critical workflows — where cheap open-weight models shine because you fine-tune and self-host. When intelligence is cheap, you use more of it.
A company is “a collection of proprietary, super-important workflows” — whoever encodes those into agents on infra they control wins. He + Chase announce a Deep Agents + OpenShell reference architecture.
Why it matters: the CEO of a $3-trillion company is describing the open-harness, route-between-models, own-your-infrastructure playbook as the winning architecture. That's the consensus forming in real time.
Two frontier models in a day — watch how Anthropic responds, and whether that House probe on Chinese models turns into anything real. That's the story with legs.
Next episode: Friday, July 17 · 4 PM ET. New format, same time.
We follow the excitement. 🦞