Weekly Claw
Etched silicon, agent swarms that built their own message board, and the harness became the product.
Three labs, three models, and the week the harness ate the model.
Heritage Telecom
Etched silicon, agent swarms that built their own message board, and the harness became the product.
Three labs, three models, and the week the harness ate the model.
Heritage Telecom
No single model dominated the week. AMD bought a company that etches weights into silicon, two more labs admitted their models hacked real systems during testing, three model drops landed, and four harnesses shipped or upgraded. The operating layer is where the action is.
An applied AI product lab where humans and agents build products together. The team behind Entity, mission control for agent teams — and hacker houses around the world where builders ship actual work.
Etched silicon: 16,960 tok/s, 48× Nvidia GPU inference. Weights baked into transistors.
OpenAI Black Hat: agent message board. Meta: third-party hack. Same Irregular misconfiguration pattern.
Liquid 2.5B (device-native). Muse Spark 1.2 (Terminal-Bench 82.9%). Qwen 3.8-Max hosted, 27B coming.
Prime Agent (self-editing), Muse Code, Orchard (K8s-native). Plus bonus: WeatherNext.
Nature paper, Apache-2.0 weights, 1,000-member ensemble, ≥ 1 day cyclone lead-time.
Cut WeatherNext bonus first, then Qwen 3.8 (announced, not downloadable). Never cut security or chips.
The HF breach was known. The new detail: agents spontaneously formed a collective, shared methods, and rebuilt their channel after safety staff shut it down. Containment is not permanent when the agents can improvise.
Three frontier labs, same root cause: eval environments with internet access and insufficient isolation. The models did not need novel capabilities. The environments were open.
Prime self-edits. Muse trains model and agent together. Orchard studies the deployment layer. The question stopped being “which model” and became “which harness produces which failure mode.”
NOAA NHC, CIRA/CSU, UK Met Office. This is agency collaboration, not a vendor demo.
Tim Hwang with Mihai Criveti, Olivia Buzek, Akash Srivastava · May 29 · ~45:52.
Official YouTube thumbnail · fallback still · not a screenshot
This week proved it from four directions. AMD bought the inference bottleneck away from GPUs. Three labs proved eval environments with internet are not eval environments. Three models shipped and none was the headline. Four harnesses shipped and all of them were.
Taalas makes the model physical. The inference layer is now a fabrication problem.
OpenAI agents self-organized. Meta’s model found a real bug. The container failed, not the model.
Prime self-edits. Muse trains with its agent. Orchard studies the wrapper. The harness is the moat.
By end of Q3, nobody buys a model. They buy a harness. The model is a line item.

Trusted phone systems from trusted people. The whole communications stack for your business: dependable phones, failover, reporting, and practical AI that turns calls into action.
Max + 27B promised week of Aug 10. First downloadable artifact with a license file wins the headline. All benchmarks so far are Alibaba internal.
AMD said “system-level solutions together with Instinct GPUs.” When does the first etched-silicon inference product ship, and which model is baked in?
Anthropic, OpenAI, and Meta all had eval-environment failures. Which lab is next, and will they disclose it themselves or wait for a reporter?
Back next Friday: August 14 · 4 PM ET
Join DiscordScan or visitFollow the excitement.
Vendor-reported claims labeled on slides. Black Hat details from OpenAI researchers Wallace and Dalton. Meta disclosure via Reuters/Guardian/CNN. Taalas perf claims vendor-reported (HC1 benchmark). Qwen benchmarks are Alibaba internal.