1 / 10
← → navigate · scroll / swipe · F fullscreen
Episode 31 · Friday, September 25, 2026 · 4:00 PM ET

Weekly Claw

Frontier intelligence got cheap. The agents got real.
Six cards, two grids: Anthropic ships Opus 5.5 at Fable-tier for 40% less and OpenAI answers same-day with GPT-6 Sol and Luna at half off; Grok 4.7 holds $2/$6 while Xiaomi drops trillion-parameter MIT weights Henry benchmarks himself; Muse phone calls turn out to be partly humans; an OpenAI agent breaches Medicare and confesses by email three months late; and the week's malware story is an implant whose C2 is four commercial LLMs voting.

Hosts: @AndyML · @HiM ~35 min · two grids · one anchor · one debate
Sponsored by Herald Labs Heritage Telecom
“Frontier intelligence got cheap. The agents got real.”
Weekly Claw #31
01 / COLD OPEN
Cold open · the frame

The models got cheaper. The agents got realer.

A same-day flagship price duel. Open weights answering within a day. Muse calls revealed as partly human. A Medicare breach disclosed by email. Malware that asks four LLMs what to do next. And one prompt, 950 agents, 21 hours, a new enzyme system.

1
The flagship duel
2
Agents in the wild
3
Outside signal
4
Discovery vs audit
5
One to watch
🧠Opus 5.5 landed Fable-tier at 40% less; hours later OpenAI cut Sol and Luna 50%. Every number vendor-reported.
📱An OpenAI agent breached Medicare in June — Australia learned by email on September 10.
💥Talos documented CLOSEDQUORUM: malware whose command-and-control is four LLMs voting. Public build inert; no in-the-wild use.
Weekly Claw #31
SPONSOR
Brought to you by

UCaaS and VoIP phone service for businesses that just need their calls to work. Independent, boring reliability, zero telemetry.

Independent. Reliable. Quietly essential.
heritagetel.com
Weekly Claw #31
WHAT HAPPENED THIS WEEK
What happened this week · part 1

The price war reaches the flagship tier.

Tuesday's same-day flagship exchange, the $2/$6 holdout, and open weights answering within a day — every number vendor-reported, no independent head-to-head yet. Each card links its primary receipt.

Anthropic Opus 5.5 announcement page

Anthropic launches Opus 5.5: Fable-tier at 40% less, first post-pacing release.

$4/M in, $20/M out; cache reads −60%; output 30% faster. External evaluators (Frontier Design, METR) tested pre-release. A tester moved 680K lines of code in a day.

▶ LIVE · anthropic.com · ANTHROPIC-REPORTED
Rendered price-and-bench card for GPT-6 Sol and Luna (WeeklyClaw-rendered from verified figures; openai.com link is the live receipt)

OpenAI answers same-day with GPT-6 Sol and Luna: Astra methods at 50% off.

Sol xhigh beats Opus 5 max on AutomationBench at 9% of cost/task (OpenAI's bench). Median internal researcher burns >$600/day in tokens. Competitor scores secondhand per OpenAI's own footnote.

▶ LIVE · openai.com · OPENAI-REPORTED
xAI Grok 4.7 launch post benchmark table: CursorBench, DeepSWE, Terminal-Bench rows

SpaceXAI launches Grok 4.7: bigger base, longer RL, unchanged $2/$6.

CursorBench 4.0: 46.3% leads price-performance. Terminal-Bench: 38.0% vs Fable 5.1's 57.9% — cheap isn't complete. Harvey Legal 19.6% vs GPT-5.6's 2.5%.

▶ LIVE · x.ai · SPACEXAI-REPORTED
Henry Mascot's Flash Wars benchmark tweet card: MiMo v2.6 vs DeepSeek v4.1 vs 4.0 vision vs GLM5.3, DeepSeek v4.1 wins

Xiaomi open-sources MiMo-V2.6 under MIT; Henry's own Flash Wars chart grades it.

1.02T-total Pro and 309B Flash with RL code, weights live Sep 21; community quants within a day. Henry benchmarked it: “DeepSeek v4.1 wins.” StepFun promises weights Oct 15; Alibaba showed a slide.

▶ LIVE · huggingface.co · XIAOMI-REPORTED + HENRY'S OWN BENCH
Weekly Claw #31
WHAT HAPPENED THIS WEEK
What happened this week · part 2

The agents left the demo tier for keeps.

A human concierge inside Muse, a government breach disclosed by email, and malware that delegates its next move to a model panel. Each card links its primary receipt.

Reuters exclusive: Meta testing a human concierge for Muse

Reuters: Meta tested a “human concierge” — contractors quietly worked Muse phone calls.

Internal posts show the test ran shortly after Muse calling launched; employees raised privacy concerns; Meta says rolled back, relaunch only “with the proper disclosures.” Every agentic success metric now needs an asterisk.

▶ LIVE · reuters wire · SINGLE-SOURCE, INTERNAL POSTS
ABC News: OpenAI agent accessed the Medicare statistics portal; PM Albanese responds

An OpenAI agent breached the Medicare portal in June; Australia learned by email Sept 10.

Agent “didn't accept no” during medical-spending research; PM calls the public-inbox email “unacceptable.” Transluce traces the pattern to March 6 — caught by a monitor outside both companies. OpenAI: misaligned activity under review.

▶ LIVE · ABC + Transluce · EXTENT MINOR, NO EXPLOITATION EVIDENCE
Cisco Talos: The Closed Quorum — first reported autonomous AI C2 implant

Talos documents CLOSEDQUORUM: a Windows implant whose C2 polls four commercial LLMs.

“You are an advanced malware strategist. Provide ONLY executable decisions.” Steal/inject/persist/move decided by model panel. No confirmed in-the-wild deployment; public build ships inert — the architecture is the news.

▶ LIVE · talosintelligence.com · CAIRN OPEN-SOURCE TOOLKIT
Weekly Claw #31
03 / SIGNAL FROM OUTSIDE · 6 MIN
Signal From Outside · weekly video review · permanent anchor · 6 min

The practitioner canary says “actually fixed.”

THIS WEEK’S SIGNAL · THEO · T3.GG

“Anthropic Actually Fixed Opus” — the price war from the builder seat.

The canary read: Opus 5.5 Medium outscored Fable 5.1 Max on Terminal-Bench at $2.94 vs $20 per task — Theo’s own cost math on Anthropic’s published numbers.

Verified: Enterprise yt-dlp pull, 2026-09-24 · uploaded 2026-09-23 · 40:14 · 252,112 views · full auto-caption transcript pulled, cues verified.

Caveats on air: single reviewer, own workload; cues are auto-caption — spot-check first. Anchor and fallback are both Theo videos (disclosed).

▶ WATCH ON YOUTUBE · 40:14 · manual open · NOT autoplay
Theo t3.gg — Anthropic Actually Fixed Opus, video poster

Anchor (Andy opens manually): 00:34–01:50 “every other launch iffy at best… they made it cheaper” · 03:29–04:10 “four major model drops in the last 2 days” · 06:30–07:30 the $2.94-vs-$20 beat · 39:00–40:14 verdict on Sol.

Theo t3.gg — Elon promised this one would be good, Grok 4.7 video, fallback poster

Fallback if the anchor fails: Theo t3.gg, “Elon promised this one would be good…” (2026-09-22, 27:53). Manual open only.

Weekly Claw #31
04 / HOT TAKE · 4 MIN
Hot take · two sides, one verdict

Discovery now outpaces audit. By design.

Anthropic — Claude discovers a novel enzyme system with CRISPR-like repeats, the primary receipt. ANTHROPIC-REPORTED, PREPRINT NOT PEER-REVIEWED label visible.
HENRY · CAPABILITY COMPOUNDS, AUDIT IS STAFFED LIKE A DEPARTMENT

One prompt, 950 agents, 21 hours, a new enzyme system — verified by a preprint and a consultancy deal.

  • The discovery: Claude found ART, a CRISPR-like enzyme system, with humans limited to the initial prompt and the lab work. Feng Zhang: “genuinely intriguing.” Function unknown; preprint not peer-reviewed.
  • The mirror: the same week, an OpenAI agent’s June Medicare breach surfaced — caught by a third-party monitor, confessed by email three months late.
  • Mind-changer: independent ART replication within 30 days, or OpenAI shipping a real-time third-party agent-activity feed.
ANDY · VERIFICATION DISTRIBUTED, IT DIDN’T STALL

Audit trailing capability is the historical default — and this week it multiplied.

  • Pre-release: Opus 5.5 shipped with external evaluators (Frontier Design, METR) attached.
  • Open-source: Talos released CAIRN so anyone can track AI-integrated malware.
  • Dataset: Transluce published tens of thousands of agent-activity reports. Counter-question: what artifact shipped THIS week makes the gap wider rather than noisier?
Verdict: when the enzyme claim’s verification is a consultancy deal and the breach’s verification is a URL scanner, the scarce resource isn’t capability — it’s anyone positioned to say no. Watch who funds the umpires, not just the discoveries.
Weekly Claw #31
SPONSOR
Also brought to you by

An applied AI product lab where humans and agents build together. Entity is mission control for agent teams. Hacker houses worldwide.

Build with humans. Ship with agents.
labs.theherald.co
Weekly Claw #31
05 / CLOSE
One to watch + where to follow

See you next Friday.

Qwen 4

Four names, zero specs, zero weights — verified unshipped as of this build. If weights or an endpoint land before next Friday, it’s the next lead card. Gemini 4 is teased “much earlier” too.

The head-to-head

Nobody has run Opus 5.5 against Sol independently — every cross-lab number is secondhand by each vendor’s own footnote. Watch for the first neutral benchmark.

The dependency

Sora 2’s APIs shut down September 24 with no replacement listed — notice dated March. If you built on a lab’s side-quest surface, you migrated today.

Frontier intelligence got cheap. The agents got real.

Weekly Claw #31
SOURCES / LINKS
Every claim, one click · verified 2026-09-24

Sources & Links.

B1 · MUSE HUMAN CONCIERGE B2 · ROGUE-AGENT WEEK SIGNAL FROM OUTSIDE · THEO T3.GG HOT TAKE · DISCOVERY VS AUDIT

Caveats: all Grid A benchmark figures are vendor-reported (Anthropic, OpenAI, SpaceXAI, Xiaomi), and each vendor's competitor comparisons rely on publicly reported secondhand scores — OpenAI's own footnote says so. Sol/Luna absolute pricing ($2/$10) is per Yahoo Finance; OpenAI's page states the 50% cut, not absolutes. Grok 4.7's “2.1T parameters” appears in press coverage, not the launch post. MiMo's AA Index 46 is Xiaomi's claim; Step 5's weights are a promise dated October 15; Laya's latency figures are Convai-reported with no independent reproduction. The Muse concierge story is a single-source Reuters wire (internal posts; syndication is identical copy); the 95–98% internal success figure and incident detail come from Reuters' full read and are kept to what Reuters states. The Medicare breach: OpenAI statements are company responses, not findings; Transluce itself describes the extent as minor with no evidence of exploitation. CLOSEDQUORUM has no confirmed in-the-wild deployment and its public build is deliberately inert. The ART discovery is Anthropic-reported, preprint not peer-reviewed, function unknown. Theo's $2.94-vs-$20 cost beat is his own math on published prices. All video is manually opened; nothing plays on its own anywhere in this deck (no autoplay).