1 / 10
← → navigate · scroll / swipe · F fullscreen
Episode 30 · Friday, September 18, 2026 · 4:00 PM ET

Weekly Claw

The models stopped chatting. The agents went to work.
Five cards, two grids: TypeSafe's Jev returns typed decisions instead of prose; an anonymous stealth model goes free at 256K context on OpenRouter; Qwen3.8-Omni-Flash ships 1M-context omni input while PrismML's 5.9 GB ternary 27B runs on phones; Apple starts the Siri AI beta across its operating systems; and Google opens Home MCP early access so third-party agents can run the house.

Hosts: @AndyML · @HiM ~32 min · two grids · one anchor · one debate
Sponsored by Herald Labs Heritage Telecom
“The models stopped chatting. The agents went to work.”
Weekly Claw #30
01 / COLD OPEN
Cold open · the frame

The models stopped chatting. The agents went to work.

Jev returns typed decisions instead of sentences. An anonymous stealth model goes free at 256K context with no nameplate. Apple puts an assistant across five operating systems. Google opens the front door of the home to third-party agents. And two launches put omni input in the cloud while a full agent stack fits on the phone.

1
Models learn to decide
2
How we know they work
3
Outside signal
4
Who pays the umpire
5
One to watch
🧠TypeSafe opened early access to Jev, a model that returns typed probability distributions rather than prose. Vendor-evaluated so far.
📱Apple started the Siri AI beta across iOS, iPadOS, macOS, watchOS and visionOS.
💥An anonymous stealth model, Union Alpha, is free on OpenRouter at 256K context — no lab, no model card. Routing receipt real, identity not.
Weekly Claw #30
SPONSOR
Brought to you by

An applied AI product lab where humans and agents build together. Entity is mission control for agent teams. Hacker houses worldwide.

Build with humans. Ship with agents.
labs.theherald.co
Weekly Claw #30
WHAT HAPPENED THIS WEEK
What happened this week · part 1

The model stops talking and starts deciding.

Three receipts about capability without a nameplate — a typed-judgment model, an anonymous 256K stealth model free on OpenRouter, and two model launches that put omni input on the cloud and a full agent stack on the phone. Each card links to its primary receipt.

TypeSafe Jev latency and cost Pareto chart from the launch post

TypeSafe opens early access to Jev, which returns typed decisions instead of text.

$40M seed out of stealth. Choice, Score and yes/no outputs at $0.042 per million input tokens, 70–500 ms claimed. Gains measured on four staff workflows.

typesafe.ai · VENDOR EVAL · 0% = schema validity, not 0% wrong
OpenRouter listing for the anonymous stealth model Union Alpha

An anonymous stealth model, Union Alpha, is free on OpenRouter with a 256K agentic target.

262K context, 131K max output; text and image in, tool calls and JSON out; also routed via OpenCode Zen with zero retention. About 2B tokens on day one, 100B+ claimed two days later. No lab, no model card, no benchmark.

openrouter.ai · ANONYMOUS PROVIDER · routing receipt real, identity not
Qwen3.8-Omni-Flash production overview and PrismML Bonsai 2 27B benchmark table

Qwen3.8-Omni-Flash ships 1M-context omni input; PrismML answers with a 5.9 GB ternary 27B that runs on phones.

Qwen: text, image, audio, video in with tool use, audio-visual costs cut over 90% (claimed), text-only output. PrismML Bonsai 2: 1.76 bits/weight, 262K context, Apache-2.0 GGUF and MLX with custom CUDA and Apple kernels; 1.8% claimed aggregate loss.

qwen.ai + prismml.com · VENDOR-REPORTED · no independent reproduction
Weekly Claw #30
WHAT HAPPENED THIS WEEK
What happened this week · part 2

The agent left the lab and met the real world.

The assistant ships to five operating systems, and Google opens the front door of the home to third-party agents. Each card links to its primary receipt.

Apple Siri AI app actions feature image from Apple's newsroom

Apple starts the Siri AI beta across its operating systems.

English beta on iOS, iPadOS, macOS, watchOS and visionOS 27. Personal context, onscreen awareness, cross-app actions. Other languages planned for October.

apple.com · BETA · availability verified, quality untested here
Google Home MCP early access announcement

Google opens Home MCP early access: third-party agents can run the house.

Named launch agents include Claude, Hermes, Open Claw and Antigravity. US-only English, Home Premium Advanced at $20/month, agents cannot unlock doors, automations not yet supported.

developers.home.google.com · EARLY ACCESS · US-only, subscription-gated, no door locks
W WEEKLY CLAW #30
03 / SIGNAL FROM OUTSIDE · 9 MIN
Signal from outside · the boring good news

The fear is loud. The good news is boring.

Jensen Huang · All-In Summit · Sep 14
Jensen Huang on stage at the All-In Summit All-In Summit · Los Angeles
“We need more radiologists than ever in the world.”
— Jensen Huang, CEO, Nvidia

Safety is paramount, he said. But safety versus leadership is a false choice. And the scary forecasts should be scored against what actually happened.

Jensen's scorecard · radiology
Predicted
AI takes over radiology. No radiologists left in five years.
→
What happened
AI reads the scans. We need more radiologists than ever.
The work changed · the people stayed · patients get read faster  ·  Jensen's claim
From outside AI · Morgan Housel · Sep 11
Psychology of Money with Morgan Housel, episode 'AI Optimism and the Agony of Waiting'
When you dread something, your mind is “a very proficient storyteller.”
— Morgan Housel, The Psychology of Money
20–30%
More productive, then life goes on.
The one forecast he'll make.
W WEEKLY CLAW #30
03 / SIGNAL FROM OUTSIDE · 9 MIN
Signal from outside · already working

A narrow model. A person makes the call.

Brain-computer interface connectors on Casey Harrell's head, UC Davis UCTV · UC Davis · Sep 8

A voice back, at home, unattended.

3,800+
hours of home use
99% word accuracy · 56 wpm
“Casey's the ultimate power user.”— UC Davis researcher
Neurons→Decoder→Voice→Casey
Ryan Honary carrying a SensoRy AI wildfire sensor in the hills OpenAI-produced · Sep 14

A teenager's sensors catch fires at ignition.

5th grade
science project →
ridge-line sensor network
“Why do you believe this is a fire?”— asked over Ryan's walkie-talkie
Sensors→LLM→Radio→Firefighter
Peter Battaglia and Hannah Fry on the Google DeepMind podcast DeepMind's own account · Sep 9

Three days' warning on a Category 5.

~3 days
lead on Hurricane Melissa's
Cat 5 call (Oct 2025)
The model's confidence raised the forecasters' own.— Peter Battaglia, DeepMind (paraphrased)
Weather data→Model→Forecast→NHC
THE PATTERN INPUT NARROW MODEL PLAIN OUTPUT HUMAN MAKES THE CALL
Weekly Claw #30
04 / HOT TAKE · 4 MIN
Hot take · two sides, one verdict

Safety needs an umpire. Who pays the umpire?

Dario Amodei — We Must Pace the Frontier, the primary essay receipt. COMMITMENT — NOT YET A CONTROL label visible.
HENRY · A COMMITMENT IS NOT A CONTROL

You don’t get to pick your own umpire and call it oversight.

  • What was said: Amodei asked for slower capability gains and permanent independent evaluators with employee-like access; Anthropic committed to the first step; Altman said OpenAI would adopt the same commitment.
  • What is missing: no evaluator named, no final access scope published, no start date — all three cells read NOT ANNOUNCED.
  • Mind-changer: a named evaluator, a published scope, a start date, and a failure it is allowed to publish.
ANDY · THE ALTERNATIVE IS SLOWER, NOT CLEANER

Name the three conditions before calling it theater.

  • Access: the essay commits evaluators to tools and internal risk-assessment processes. That is more than a PDF review.
  • Publication: oversight without a public dissent route is a memo. Publication rights are the load-bearing clause.
  • Counterfactual: state regulators move slower and carry their own politics. Embedded scrutiny at least produces dated artifacts.
Verdict: an embedded evaluator is the strongest oversight proposal these labs have put in writing. It becomes oversight when the evaluator can name a lab failure, publish it, and still hold the badge afterwards. Until then the commitment is a promise — and the promise is inspectable.
Weekly Claw #30
SPONSOR
Also brought to you by

UCaaS and VoIP phone service for businesses that just need their calls to work. Independent, boring reliability, zero telemetry.

Independent. Reliable. Quietly essential.
heritagetel.com
Weekly Claw #30
05 / CLOSE
One to watch + where to follow

See you next Friday.

The second reporter

OpenAI’s Ready-for-Disclosure track now has a precedent. Watch for the first case that is not OpenAI reporting on OpenAI: a dated third-party incident report naming the control that was missing.

The repeat test

Watch who attaches a nameplate to Union Alpha — or who quietly routes production work through it while it stays anonymous. The routing receipt is real; provenance is not.

The verification

Home MCP names Claude, Hermes, Open Claw and Antigravity as launch agents but blocks door locks and automations. Watch for the first door, alarm or lock endpoint — and who audits it.

The models stopped chatting. The agents went to work.

Weekly Claw #30
SOURCES / LINKS
Every claim, one click · verified 2026-09-18

Sources & Links.

B1 · APPLE SIRI AI BETA B2 · GOOGLE HOME MCP SIGNAL · BORING GOOD NEWS + ALREADY WORKING HOT TAKE · WHO PAYS THE UMPIRE

Caveats: Jev’s speed, cost and calibration figures are TypeSafe-evaluated, the 193.6×/444.6× averages run on four staff-authored workflows with GPT-6 Astra and Claude Fable 5.1 reference labels, and the 0% hallucination plot is a schema guarantee rather than accuracy. Union Alpha has no lab, no model card and no benchmark; its context, price and routing fields are machine-verifiable on OpenRouter, its usage figures are platform-claimed, and the 36kr report is secondary. Qwen’s over-90% audio-visual cost reduction is vendor-reported and output is text-only; QwenCloud lists 991K max input against the 1M headline. PrismML’s benchmark table, throughput and 1.8% aggregate loss claim are publisher-reported with no independent reproduction; run the tool-loop eval before trusting the ternary kernel. Apple’s beta verifies availability, not quality or the privacy architecture under adversarial testing; server-backed features carry daily limits. Home MCP is US-only English, gated to Home Premium Advanced at $20/month, cannot unlock doors, and does not yet support automations. The evaluator commitments have no named evaluator, published scope, or start date. All video is manually opened; nothing plays on its own anywhere in this deck.