1 / 10
← → navigate · scroll / swipe · F fullscreen
Weekly Claw #32
EPISODE 32
Episode 32 · October 2, 2026 · 4:00 PM ET

Weekly Claw

New models and new controls for agents

Henry & Andy

Weekly Claw #32
01 / COLD OPEN

What would you change in your agent stack?

1OpenAI launches Sol, dots and a Decisions API preview
2Google introduces Argon with restricted initial access
3New runtime controls face an independent-evidence test
Weekly Claw #32
SPONSOR
Brought to you by

UCaaS and VoIP for businesses that need their calls to work

Quietly essential
heritagetel.com
Weekly Claw #32
WHAT HAPPENED THIS WEEK

OpenAI and Google launch models and agent tools

WeeklyClaw benchmark table transcribed from OpenAI: DeepSWE high +6.4 percentage points vs prior Sol best, AutomationBench medium +2.2 points vs Opus 5.5, OSWorld offline max 2.1 points below Astra max

OpenAI launches GPT-6.1 Sol

OpenAI-reported results, with effort levels retained

▶ Benchmark source
Original OpenAI dots launch film frame at 1:05 showing Jojo’s Computer browsing a wedding cake website

OpenAI launches always-on dots

Official demo; plan, market and admin limits apply

▶ Official launch video
Original OpenAI Decisions API demo at 0:10: routing 10000 customer requests, 150 ms per request versus Responses API 1.6 seconds

OpenAI previews the Decisions API

Vendor demo shown 15× realtime; limited preview

▶ Official demo video
Selected coding rows from Google’s original benchmark table: Argon 77.9 DeepSWE, 55.0 FrontierSWE, 91.9 Vibe Code, 57.4 Terminal-bench 4.0 with Astra and Claude comparison columns

Google introduces Gemini 4 Argon

Google comparison; mixed-source evaluations

▶ Full table and methodology
Weekly Claw #32
WHAT HAPPENED THIS WEEK

Agent controls face new tests and scrutiny

Source-verified benchmark scorecard showing GLM-5.3 50 of 410 end-to-end exploits and Mythos Preview 56 of 410, calculated 12.2 versus 13.7 percent

Anthropic evaluates GLM-5.3 cyber capability

Anthropic-run; capability test with safeguards disabled

▶ Evaluation and caveats
OpenClaw’s original launch illustration of three coral lobster agents in separate rooms connected to a shared workspace

OpenClaw Enterprise opens for internal pilots

Self-hosted preview; internal pilots before 1.0

▶ Official launch and repo
NVIDIA original diagram: OpenShell gateway manages three agent sandboxes and policy-approved application-service connections

NVIDIA launches an agent safety platform

OpenShell runtime; Sentry watchdog on BlueField-4

▶ Official tutorial and architecture
Excerpt of METR’s testimony cover identifying the rogue-AI hearing and September 30, 2026 date

Regulators and lawmakers scrutinize AI agents

FTC investigation; Senate testimony; lawsuit allegations

▶ Testimony and source links
Weekly Claw #32
03 / SIGNAL FROM OUTSIDE
Jensen Huang · The Ezra Klein Show · September 23

Jensen and Ezra debate AI safety

Official thumbnail for Jensen Huang’s interview with Ezra Klein
Jensen Huang Thinks A.I. Alarmism Has Gone Too Far · 1:47:21

Safe enough to ship?

What evidence should stop a release?

NVIDIA’s CEO makes the case.
Ezra Klein challenges it.

Weekly Claw #32
04 / HOT TAKE
The motion

Agent vendors should be required
to publish incident reports

FOR

Customers need evidence to compare deployment risks

A disclosure deadline creates accountability

Hard question: Who verifies completeness?

AGAINST

Disclosure can expose vulnerabilities and private data

Reporting rules can punish smaller vendors

Hard question: What replaces public disclosure?

What evidence would change your position?

Weekly Claw #32
SPONSOR
Also brought to you by

An applied AI product lab where humans and agents build together

Entity is mission control for agent teams

Build with humans. Ship with agents.
labs.theherald.co
Weekly Claw #32
05 / CLOSE
One to watch

Can another team reproduce the safety result?

A repeatable boundary test, its failure cases,
and the resulting audit trail

The evidence we want to see before expanding a pilot

Weekly Claw #32
SOURCES / LINKS