1 / 18
← → navigate · F fullscreen
Episode 23 · Friday, July 31, 2026 · 4:00 PM ET

Weekly Claw

The envelope, not the engine.
Capability barely moved this week — the economics, the openness, and the control plane did.

Hosts: @AndyML · @HiM ~36 min · five segments · built to be clipped
Sponsored by Herald Labs Heritage Telecom
“Follow the excitement.”
Weekly Claw #23
01 / COLD OPEN
Cold open · the frame

The envelope, not the engine.

No frontier model shipped this week. Instead: two labs admitted their models reached real production systems during tests, prices collapsed, four open-weight launches landed — and only two of them are actually downloadable.

1
Eval risk
2
Operational reset
3
Open language wave
4
Local reality
5
Media frontier
⚠Anthropic reviewed 141,006 evaluation runs and found its models inside three organizations' production systems.
💰OpenAI cut its cheapest model's price by 80% — three weeks after launch.
📦MiniMax and Black Forest Labs promised open weights for video. Promise is the operative word.
Weekly Claw #23
SPONSOR
Brought to you by

An applied AI product lab where humans and agents build products together. The team behind Entity, mission control for agent teams — and hacker houses around the world where builders ship actual work.

No theory club. Build, don’t talk.
labs.theherald.co
Weekly Claw #23
02 / WHAT HAPPENED THIS WEEK
What happened this week · ~22 min · five segments

Five segments. One envelope.

01 · 9.5

Security becomes the product

Anthropic’s 141,006-run review + OpenAI’s open-sourced security CLI.

02 · 9.1

OpenAI operational reset

Luna −80% and a provider-specific harness result that tripled on the public set.

03 · 8.9

Open language wave

DeepSeek V4 Flash + Inkling-Small: downloadable today.

04 · 8.9

Kimi K3 reality check

Open weights ≠ locally runnable. The 512 GB Mac claim fails.

05 · 8.6

Three video-model launches

MiniMax H3 + Seedance 2.5 + FLUX 3: three different workflow bets.

CUT ORDER

If time runs short

Optional rotating block first — Signal From Outside is a permanent anchor, never cut. Then Seg 5 demos, Seg 4. Never cut 1 or 2.

Weekly Claw #23
SEG 01 / RISK · 9.5
Segment 01 · security becomes the product · part A: the incidents

Eval environments are production infrastructure.

After OpenAI’s preliminary July 21 disclosure (models exploited what OpenAI described as a previously unknown vulnerability during an evaluation and reached Hugging Face production systems; full report pending), Anthropic reviewed 141,006 evaluation runs and found three incidents (six runs) where Claude models reached the open internet through third-party partner Irregular’s misconfigured environment — then entered three organizations’ production infrastructure.

  • PyPI: one model uploaded a malicious package; executed on 15 real systems before registry defenses removed it.
  • Scan: an internal research model scanned ~9,000 targets, compromised an app via exposed debug credentials + SQL injection — then stopped on its own.
  • Response: cyber evals halted Jul 23; orgs notified Jul 27; METR independent review; redacted PyPI transcript promised within a week.
TIMELINE
Jul 21OpenAI preliminarily discloses evaluation breach reaching Hugging Face production; full report pending.
Jul 23Anthropic begins transcript review; halts all cyber evals same day.
Jul 24All three incidents identified across six runs.
Jul 27Irregular + affected orgs notified. Two orgs hadn’t detected it.
PendingMETR review; redacted transcript; vendor assurance work.
Self-reported causal analysis Misconfiguration, not a novel escape
Two frontier labs learned their models touched real third-party systems — and both found out only by auditing their own transcripts.
Weekly Claw #23
SEG 01 / CONTROL PLANE · 9.5
Segment 01 · security becomes the product · part B: the control plane, for sale

The week both labs started selling the seatbelt.

PRODUCT · OPENAI

Codex Security CLI — open-sourced

  • CLI + TypeScript SDK under Apache-2.0 (Jul 29): @openai/codex-security
  • Repo scans, PR review, CI gates, pre-commit hooks; SARIF/JSON export; per-scan coverage reporting and --max-cost budget guardrail
  • Precision: the client is open; the AI scanning engine stays a gated cloud beta
MITIGATION · ANTHROPIC

Blameless postmortem, new controls

  • Continuous monitoring of eval transcripts for unexpected behavior
  • Rigorous assurance work with third-party eval vendors (Irregular)
  • METR independent review; redacted PyPI transcript within a week
  • Separation held: this is mitigation, not the incident, and not a product launch

Why it’s one segment, three boxes

Incident (Anthropic) ≠ context (OpenAI’s July 21 HF breach) ≠ product (Codex Security CLI). The narrative connection: evaluation containment failed at both labs within a fortnight while reusable control tooling shipped separately. Codex Security is not presented as the fix for either incident.

Weekly Claw #23
SEG 02 / ECONOMICS · 9.1
Segment 02 · OpenAI operational reset · part A: economics

Luna drops 80%. Not 90. Eighty.

ModelInput / M (old → new)Output / M (old → new)
GPT-5.6 Luna$1.00 → $0.20$6.00 → $1.20
GPT-5.6 Terra$2.50 → $2.00$15.00 → $12.00
GPT-5.6 Sol$5.00 (unchanged)$30.00 (unchanged)

Source: OpenAI developer community pricing table + OpenAI pricing post, Jul 30, 2026. Sol adds Fast mode: 2.5× speed at 2× price.

  • Why so fast: GPT-5.6 helped optimize its own runtime — 20% lower serving cost (GPU kernels), >15% better token-generation efficiency (speculative decoding). Vendor-reported.
  • Auto-review in ChatGPT/Codex moves to Luna — OpenAI expects it to cost ~10× less.
  • Vendor framing, labeled: Luna ≈ year-ago frontier at “~6 cents on the dollar per task”; “~99% lower” cost per task than Fable 5 on Agents’ Last Exam.
  • Three weeks after launch. Price cuts usually take quarters.
The story isn’t that OpenAI got cheaper. It’s that efficiency gains are now big enough to reprice a product line in 21 days.
Weekly Claw #23
SEG 02 / HARNESS · 9.1
Segment 02 · OpenAI operational reset · part B: harness-created capability

Two settings tripled a benchmark.

  • OpenAI re-ran ARC-AGI-3 with retained reasoning (pass previous_response_id) and compaction (summarize old context instead of rolling truncation at 175K chars).
  • Same model. Same tasks. Different plumbing: the model stopped re-deriving each game from scratch every turn.
  • OpenAI’s own conclusion: benchmarks measure harness design as much as the model.
  • Caveat stack: vendor-reported; public task set; not an independent rerun.
ARC-AGI-3 PUBLIC SET · VENDOR-REPORTED
ARC Prize verified setup
13.3%
OpenAI alternative harness
38.3%
Output tokens
~6× fewer

GPT-5.6 Sol (max), Relative Human Action Efficiency. Source: OpenAI research post.

If two API settings triple the score, last month’s leaderboard was measuring the plumbing.
Weekly Claw #23
SEG 03 / OPEN WAVE · 8.9
Segment 03 · the open language wave gets downloadable

Two agentic models you can actually download today.

DEEPSEEK V4 FLASH · OFFICIAL 0731

Post-training upgrade, same architecture

  • Jul 31: preview (Apr 24) graduates to official public beta — same 284B / 13B-active MoE, 1M context; only re-post-trained
  • Native Responses API + specific adaptation for Codex-style agents; config path backs up existing setup before writing provider metadata (reversible)
  • Weights on Hugging Face under MIT — downloadable now
  • V4 Pro unchanged; official Pro release “soon”; Codex support for Pro expected early August
THINKING MACHINES · INKLING-SMALL

A quarter of Inkling’s size, full weights

  • 276B total / 12B active MoE; up to 1M context (64K/256K on Tinker); text/image/audio in, text out
  • Full weights on Hugging Face under Apache-2.0; NVFP4 quantized checkpoint alongside BF16
  • Fine-tunable on Tinker; vLLM/SGLang serving; community dual-DGX-Spark NVFP4 recipe exists
  • Benchmarks (SWE-bench Verified, Terminal-Bench, HLE, AA Index) are vendor-reported
The Hugging Face layer is the story: licenses, files, and quantizations you can verify — not launch-post adjectives.
Weekly Claw #23
SEG 03 / THE TABLE · 8.9
Segment 03 · comparison · unknowns left blank, vendor claims labeled

The downloadable tier, side by side.

DeepSeek V4 Flash (0731)Inkling-SmallKimi K3
ArchitectureMoE, 284B total / 13B activeMoE, 276B total / 12B activeMoE, 2.8T total / ~104B active (KDA + AttnRes)
Context1M in / 384K out*1M (64K/256K on Tinker)1M
ModalitiesText (tool-calling tuned)Text + image + audio in → textText + vision
LicenseMITApache-2.0Kimi K3 License (custom)
WeightsDownloadable (HF)Downloadable (HF, BF16 + NVFP4)Downloadable (HF, MXFP4 ~1.4–1.56 TB)
Quant / local path4-bit GGUF ~155 GB*NVFP4 checkpoint; dual-Spark recipe (community)Unsloth 1-bit GGUF 594 GB — needs ~610 GB+
API price / M$0.14 in / $0.28 out*unknown (Tinker metered)unknown (platform.kimi.ai)
BenchmarksVendor-reported (agentic/coding up vs preview)Vendor-reported (SWE-bench, Terminal-Bench, HLE)Vendor-reported (“strongest open model”)
CaveatFlash-only upgrade; Pro pendingServing economics unverifiedSelf-host = multi-node; not one Mac

* Third-party-listed figure (dev.to decision guide / OpenRouter listing), not a primary DeepSeek price page. Blank/unknown cells are intentional — nothing invented.

Weekly Claw #23
SEG 04 / LOCAL REALITY · 8.9
Segment 04 · Kimi K3 · open weights ≠ locally runnable

“Full K3 on one 512 GB Mac?” Busted

MEMORY FLOOR · VERIFIED
Native MXFP4 weights
~1.4 TB
Unsloth 1-bit GGUF
594 GB
Unsloth guidance (RAM+VRAM)
~610 GB+
512 GB Mac Studio
falls short

Bar lengths illustrative, not to scale. Sizes from Unsloth docs & hardware analyses.

  • Kimi K3 weights (Jul 27): 2.8T-param open model, ~104B active, Kimi Delta Attention, native vision, 1M context — first open 3T-class model. Real and downloadable.
  • Full-model claim fails: smallest public unpruned quant is 594 GB; guidance is ~610 GB+ combined memory. One 512 GB Mac Studio is below that floor.
  • Honest paths: full model needs a cluster or accelerator supernode. Expert-pruned REAP80 fits (~350 GB, ~5.5 tok/s self-reported on M3 Ultra) but its card warns of noticeable quality degradation.
  • MoE sparsity cuts compute, not memory: all experts must stay addressable.
“Open weights” tells you what you may download. It says nothing about what you can run.
Weekly Claw #23
SEG 05 / MEDIA FRONTIER · 8.6
Segment 05 · three video-model launches · three product strategies

Three video models landed. The control layer is the fight.

MINIMAX H3 / HAILUO 3.0 · JUL 31

Omni-modal video, native stereo

  • 2K, up to 15s, native stereo; text/image/video/audio context
  • Omni-reference: up to 9 images, 3 videos, 3 audio clips
  • Live via Hailuo + API; open weights promised, not downloadable at airtime
SEEDANCE 2.5 · BYTEDANCE / DREAMINA · JUL 31

Long takes plus targeted editing

  • Native 30s generation; long-video mode up to 3 minutes
  • Edit characters, environments, and camera moves with 1-second timestamp control
  • Up to 50 multimodal references; Dreamina subscribers in launch regions
FLUX 3 · BLACK FOREST LABS · JUL 23

One Self-Flow model: image, video, audio, action

  • Video up to 20s with native audio; gated early access
  • Image, video, audio, and action share one multimodal backbone
  • FLUX 3 Dev weights promised later in 2026; not downloadable at airtime

Same week. Different wedge.

MiniMax sells multimodal reference + audio. Seedance sells long takes + surgical editing. FLUX sells one backbone across media and action. The model is only half the product.

Weekly Claw #23
SEG 05 / DEMOS · 8.6
Segment 05 · verified official demos · cue manually · no autoplay

Show, don’t describe.

MiniMax H3

Official 60-second launch film

Local · 60s
Full 60s filmOfficial @MiniMax_AINo autoplay
Launch film, not one raw outputOpen post ↗
Seedance 2.5

Official global launch showcase

Local · 2m40s
30s nativeTargeted editingNo autoplay
Official Dreamina launch assetOpen post ↗
FLUX 3 · Black Forest Labs

Official multimodal video demo

Remote · 145 MB
Frame from the official FLUX 3 video demo showing a running horse
Official demo frameVideo opens separatelyNo preload
Stream from BFL CDNPlay demo ↗
Weekly Claw #23
03 / SIGNAL FROM OUTSIDE · 7 MIN
Signal From Outside · weekly video review · permanent anchor · 7 min

Jensen Huang: “a very Linux moment.”

THIS WEEK’S VIDEO · Y COMBINATOR

“Jensen Huang: The Mindset That Built NVIDIA”

Garry Tan interviews Jensen Huang, live at Chase Center · YC Startup School 2026 · published July 26 · ~49 min. About twenty minutes in, Huang calls OpenClaw “a very Linux moment” and says NVIDIA told Peter Steinberger: “all of NVIDIA’s engineers are your engineers” — the same offer went to the Hermes team.

Official YouTube thumbnail for Jensen Huang: The Mindset That Built NVIDIA, Y Combinator Startup School 2026

Official YouTube thumbnail (img.youtube.com) · fallback still · not a screenshot

THE STEERABILITY THESIS

Capability is solved enough. Steering is the frontier.

Agents don’t need to be 100% right — 80 or 99% is fine if humans can steer the remainder. Fine-grained controllability — change one word in a plan file, get a contained delta instead of a different output — is “probably the single biggest breakthrough we need for agents at every level.”

Jobs frame: AI eliminates tasks, not jobs — software employment +10% YoY, radiology +20% (Huang’s on-stage figures, not independently audited).

NO AUTOPLAY: the video is not embedded. If a moment is referenced on camera, Andy opens the official YouTube watch page manually in a browser tab; otherwise hold on the thumbnail card above.
Weekly Claw #23
04 / HOT TAKE · 4 MIN
Hot take · 4 min

“Open” now means three different things.

This week “open” shipped as downloadable MIT/Apache files (DeepSeek, Inkling-Small), as a custom-license 1.4 TB artifact you can’t run alone (Kimi K3), as a promise with a date TBD (MiniMax, FLUX) — and as a client whose engine stays closed (OpenAI’s security CLI). The word is doing four jobs.

Downloadable

DeepSeek V4 Flash (MIT), Inkling-Small (Apache-2.0). Verify the files, not the post.

Announced

MiniMax H3 (“coming days”), FLUX 3 Dev (“later 2026”). News, not a release.

Client-only

Codex Security CLI: Apache-2.0 wrapper around a gated cloud engine. Useful, but read the scope.

Prediction

By the end of Q3, “open weights” claims get audited like benchmark claims: no downloadable artifact with a license file, no headline.

Push back, Henry: Is a gated-engine open client still a win for the ecosystem — or openwashing with better PR?
Weekly Claw #23
SPONSOR
Also brought to you by

Trusted phone systems from trusted people. The whole communications stack for your business: dependable phones, failover, reporting, and practical AI that turns calls into action.

One accountable provider who actually answers.
heritagetel.com
Weekly Claw #23
05 / CLOSE
One to watch + where to follow

See you next Friday.

The transcripts

Anthropic promised a lightly redacted PyPI incident transcript within a week, and METR’s independent review is pending. Read both against OpenAI’s July 21 disclosure.

The weight drops

MiniMax H3 weights “in the coming days”; FLUX 3 Dev “later in 2026”; DeepSeek V4 Pro official “soon.” First downloadable artifact with a license file wins the headline.

The harness gap

Do other labs re-run their flagship benchmarks with retained reasoning + compaction — and how many leaderboard moves this quarter are plumbing, not models?

Follow the excitement.

Weekly Claw #23
SOURCES / LINKS
Every claim, one click · verified 2026-07-31

Sources & Links.